Blog Field Notes · POV

Many Small AI Models vs One Large: We Run 420+ Models in Production. Here's Why Bigger Isn't the Answer

Many Small AI Models vs One Large: The Numbers Speak

Running 420+ models behind a single gateway beats one monolithic model every time — because in India's fitness market, specialization at scale wins over generalization at scale. We serve 65,870 customers across 11 lines of business and roughly 365 gyms, and no single model — no matter how large — can handle that breadth without bleeding money and latency. The debate of many small ai models vs one large is not academic for us. It's a daily operational reality measured in rupees, milliseconds, and Reels shipped.

A Monolith Can't Hit ≤800ms Voice-to-Voice

Our voice agents handle 10,000+ AI voice calls a day at about ₹6–10 per call versus ₹50+ for a human. That economics only works because we route each call through the smallest, fastest model that can handle the specific task — membership renewal, class scheduling, complaint resolution, lead qualification, trainer assignment. A single giant model would burn through compute on every interaction, inflating cost and latency simultaneously. Our ≤800ms voice-to-voice benchmark is non-negotiable for India's phone-first consumers. When a Hinglish-speaking member calls about their trainer slot, the model handling that exchange is tuned for exactly that conversation — not for poetry, not for code, not for general knowledge. It's tuned for +91 consumer intent at scale. A monolith can't promise that because it wasn't built for one job — it was built for all of them, and it pays a tax on every one.

The WTF Video Engine Proves the Point

Our Video Engine is a self-orchestrating factory: one brief enters, an eight-stage pipeline runs, and a brand-correct Reel exits at about $0.30–$1.70 per finished Reel versus $80–$200 for a UGC shoot. We push roughly 100 Reels a week through this pipeline. Each stage — scripting, voiceover, visual selection, captioning, compliance, formatting — is handled by a different specialized model behind our gateway. None of them is a generalist. All of them are small, fast, and individually replaceable. The scripting model doesn't need to know about captioning. The compliance model doesn't need to know about voiceover. Each one does one thing exceptionally well, and the pipeline orchestrates them like a factory line — not like a single artisan trying to build everything by hand. When one stage needs an upgrade, we swap the model for that stage alone. The pipeline never stops. That's the power of many small ai models vs one large: surgical improvement without systemic risk.

WhatsApp at 0% BSP Markup Requires Surgical Cost Control

We run WhatsApp direct on the Meta Cloud API at 0% BSP markup. Every rupee matters when you're operating at India's price points across 11 lines of business. A single large model would burn through margin on every message, every template, every automated reply. Instead, we keep 164 templates synced and 157 approved, each routed to the model that handles it cheapest without quality loss. The gateway decides in real time which model serves which request. That's not possible when one model does everything at one price point — because that price point is always the most expensive one.

By the Numbers

  • 65,870 customers served across 11 lines of business and roughly 365 gyms — no monolith can specialize for that spread.
  • 10,000+ AI voice calls a day at ₹6–10 per call versus ₹50+ for a human — only small-model routing makes that viable.
  • ≤800ms voice-to-voice latency — a monolith can't guarantee it; our gateway does.
  • 420+ models behind one gateway — each one chosen for cost, speed, and task-fit.
  • About $0.30–$1.70 per finished Reel versus $80–$200 for a UGC shootroughly 100 Reels a week, eight stages, zero generalists.
  • 164 templates synced, 157 approved on WhatsApp at 0% BSP markup — every message routed to the cheapest sufficient model.

The Counter-Argument

The objection is predictable: one large model is easier to maintain, easier to prompt, and gets better with scale. We disagree on all three counts. Maintaining 420+ models behind one gateway is not harder — it's just different. Our gateway handles routing, fallback, versioning, and observability. When a model underperforms, we swap it without touching the rest of the system. When a monolith underperforms, you're stuck retraining, reprompting, or waiting for the next release of the entire beast. Easier to prompt? Maybe — but easier to prompt badly at ₹50+ per call and $200 per Reel. Specialization is an engineering choice, not a compromise. India's market doesn't reward good enough across everything. It rewards sharp, fast, and cheap on the thing that matters right now.

What This Means

  • India's market rewards specialization over generalization — always has, always will.
  • 420+ models behind one gateway is not a vanity metric; it's a cost-and-latency survival strategy.
  • At ₹6–10 per call and $0.30–$1.70 per Reel, the unit economics only work when each task hits the smallest sufficient model.
  • If you're betting on one giant model for India-scale operations, you're betting on the wrong cost curve.
  • Agentic systems aren't tools — they're a workforce, and workforces specialize.