Small dedicated teams building bespoke autonomous agents inside your own entity, with evaluation and observability treated as the core engineering problem rather than an afterthought.

The demo is rarely the hard part. What separates a prototype from a production agent is knowing, continuously and specifically, whether it is working.
Tool definitions, planning strategy, memory and the orchestration between steps. The design decisions here determine cost, latency and failure behaviour far more than the choice of model does.
Offline evaluation sets, regression suites and a failure taxonomy. This is the most important component and the one most commonly absent, which is why so many agents never leave the prototype stage.
Full traces of every run, with the inputs, tool calls, intermediate reasoning and outcome captured. Without this, debugging an agent in production is guesswork.
The connections into your actual systems, where most of the real engineering effort lives. An agent is only as capable as the tools it can reliably call.
Input and output validation, permission boundaries, cost ceilings and human in the loop checkpoints where the action is consequential or irreversible.
Token budgets, caching, model routing and step limits. Agentic systems fail commercially long before they fail technically, usually through unbounded loops.

An agentic squad is typically four to eight people. A lead engineer who owns architecture and evaluation strategy, two or three engineers who have taken systems into production, a data or platform engineer, and someone who owns evaluation full time.
That last role is the one companies leave out, and leaving it out is the single most reliable predictor that the system will stall between demo and deployment. Knowing whether an agent is working is a full time engineering problem, not something the team does at the end.
Because the squad sits inside your own entity, the models, prompts, evaluation sets and tooling are your intellectual property. This matters more in agentic work than in most engineering, because the evaluation sets and failure taxonomy you accumulate are the genuinely durable asset.
| Dimension | Dedicated squad (Nano GCC) | AI consultancy | Hiring in your home market |
|---|---|---|---|
| Who owns the evaluation sets | You | Usually shared | You |
| Who owns the IP | Your entity | Contract dependent | You |
| Continuity after launch | Permanent team | Ends with engagement | Permanent |
| Time to a working squad | 3 to 4 months | 2 to 4 weeks | 6 to 12 months |
| Fully loaded cost per engineer | Lowest | Highest | Highest |
| Senior AI talent availability | Deep pool in India | Vendor allocated | Severely constrained |
| Knowledge retained in house | Yes | Rarely | Yes |
| Suitable for long lived systems | Yes | No | Yes |
This distribution consistently surprises teams whose expectations were set by a prototype.
The constrained hire, brought in first, who then participates in every subsequent interview. Genuine production experience with agentic or machine learning systems is worth waiting for rather than substituting.
A specific, bounded use case chosen with a measurable success definition and an evaluation set built before implementation begins. Choosing the wrong first use case is the most common failure and the most expensive.
Engineers with production experience, a data or platform engineer, and a dedicated evaluation owner. Deliberately not hired all at once, so the lead can calibrate the bar.
A narrow agent in real use with full tracing, guardrails and a cost ceiling, rather than a broad one in a demo environment. Narrow and deployed beats broad and pending.
Additional use cases added once the evaluation harness and operational patterns are proven, reusing the infrastructure rather than rebuilding it each time.
The evaluation data and failure taxonomy accumulated over a year are the durable asset in agentic work. In a consultancy engagement these frequently walk out of the door at the end.
The engineers who debugged last quarter’s failure mode are the ones designing around it now. Agentic systems fail in specific, learnable ways, and that knowledge does not transfer through documentation.
A fixed team cost rather than a day rate that scales with scope, which matters in a field where scope genuinely does expand as you learn.
India has one of the largest pools of engineers with production machine learning experience, at a fully loaded cost that makes a dedicated squad viable rather than exceptional.

The squad is the deliverable, but what makes it effective in month six is the evaluation and observability infrastructure built in month two.
Agentic capability is intended to be a durable part of your product rather than an experiment, the evaluation sets and failure knowledge you accumulate have real value, and you want that knowledge retained in house across model generations.
The task is well specified, deterministic and high volume, where conventional automation is cheaper and more reliable. Also when you want a single proof of concept and nothing after it, which a consultancy will deliver faster and at lower commitment.
One honest caveat. Agentic AI is not the right tool for every problem, and a squad will tell you so faster than a vendor will. If the task is well specified, deterministic and high volume, conventional automation is cheaper, more reliable and easier to reason about. We would rather scope a smaller agentic surface that reaches production than a broad one that stalls in evaluation.
Four to eight people for a first production system. A lead engineer, two or three engineers with production experience, a data or platform engineer, and a dedicated evaluation owner. Smaller than four tends to mean evaluation gets skipped, which is the failure mode that matters.
Because agentic systems fail in ways that are not obvious from the output. An agent that is right eighty percent of the time and confidently wrong the rest is worse than no agent, and you cannot tell which is which without an evaluation harness. It is the difference between a demo and a deployment.
Yes. The engineers are employed by your Indian entity with IP assignment in their contracts, so the architecture, prompts, tooling and evaluation data vest in you directly. There is no vendor in the chain of title, which matters most for the evaluation sets, since those are the hardest asset to rebuild.
Token budgets per run, step limits, caching, model routing so that cheaper models handle simpler steps, and hard cost ceilings with alerting. Unbounded loops are the most common way agentic systems become commercially unviable, and the controls belong in the architecture rather than in a dashboard.
Yes, and this is the common pattern. The India squad typically owns the agentic layer, evaluation infrastructure and integration work while your existing team retains model and data ownership. Boundaries are agreed at the start rather than negotiated later.
Narrow, internally facing, with a measurable success definition and tolerance for error. Retrieval over internal documentation, triage and routing, or structured extraction are good starting points. Customer facing and irreversible actions should come after the evaluation harness has proven itself.
Four to six months from squad formation to a narrow agent in real use, assuming the entity is being set up in parallel. Anyone promising production agentic systems in six weeks is describing a prototype, which is a different thing and will not survive contact with real inputs.
It will, which is an argument for owning the squad rather than against it. Model choice is the most replaceable layer in the system. Evaluation sets, integration work and operational knowledge are the durable parts, and a permanent team carries those across model generations.
Tell us the use case you have in mind and what a good outcome would look like. We will come back with a squad shape, a realistic path to production, an evaluation approach and a fully loaded cost.