We Run 420+ Models in Production: Many Small AI Models vs One Large Is Not Even Close
Many small AI models vs one large: the fight is already over
At WTF, we have made our bet on many small AI models vs one large—and the numbers make it obvious. We run 420+ models behind one gateway in production, each tuned for a narrow job, and together they move 10,000+ AI voice calls a day across roughly 365 gyms and 11 lines of business. The idea that you need one giant, general-purpose model to do real work in India is a story sold by people who have never had to hit a unit economics target in Delhi.
The thesis is simple: specialization beats generalization when you are operating at the cost and latency floor of the Indian market. A single large model can do many things poorly. A swarm of small models, routed intelligently, does each thing at a fraction of the cost, at a fraction of the latency, with a fraction of the compute. That is not a philosophical preference. That is what our production numbers say every single day.
420+ models behind one gateway means you pay for what you use
Our gateway routes every request to the cheapest, fastest model that can do the job. When a customer calls to freeze a membership, that is a different model than the one handling a Hindi sales pitch. When our Video Engine renders a brand-correct Reel, eight different specialized models handle scripting, voiceover, visuals, captions, compliance, and assembly. You do not spin up a giant brain to tie a shoelace.
The cost math is brutal and one-directional. We run about 10,000+ AI voice calls a day at about ₹6–10 per call. A human call costs ₹50+. Our voice-to-voice latency is ≤800ms. None of that survives if you route everything through one massive model. The gateway is the product. The models are interchangeable, replaceable, and ruthlessly optimized.
The WTF Video Engine proves the thesis on the content side
Our Video Engine is a self-orchestrating AI video factory. One brief goes in. Eight stages later, a finished, brand-correct Reel comes out. Behind that pipeline are 420+ models, each handling a sliver of the work—text generation, image synthesis, voice cloning, editing, formatting for platform, brand compliance, and final QC.
The result: about $0.30–$1.70 per finished Reel versus $80–200 for a UGC shoot. We produce around 100 Reels a week. If we had tried to do this with one large model doing everything, the cost per Reel would be unsustainable, the latency would be absurd, and the brand consistency would collapse. Specialization is what makes the factory work.
India does not reward generalists
The Indian market is the most price-sensitive, language-diverse, latency-punishing market on the planet. We serve 65,870 customers across roughly 365 gyms. We run WhatsApp direct on the Meta Cloud API at 0% BSP markup with 164 templates synced and 157 approved. Every rupee matters. Every millisecond matters. Every Hinglish utterance matters.
A generalist model does not know that a customer in Lajpat Nagar speaks differently than a customer in Indiranagar. Our specialized models do—because we trained and tuned them on our data, for our context, for our 11 lines of business. That is the difference between a tool and a workforce. We do not have AI tools. We have an agentic workforce, and each agent has a job description.
- 420+ models behind one gateway, each tuned for a narrow task
- 10,000+ AI voice calls/day at ₹6–10 per call vs ₹50+ for a human
- ≤800ms voice-to-voice latency—unreachable with a single giant model
- About $0.30–$1.70 per finished Reel vs $80–200 for a UGC shoot
- Around 100 Reels/week through the Video Engine's eight-stage pipeline
- 164 templates synced, 157 approved on WhatsApp at 0% BSP markup
- 65,870 customers across roughly 365 gyms and 11 lines of business
The counter-argument: isn't managing 420+ models a nightmare?
The objection is always the same: one model is simpler to operate. One model means one vendor, one prompt strategy, one upgrade path. 420+ models means 420+ things that can break, drift, or degrade.
Our answer: that is what our gateway is for. The gateway abstracts the complexity. Our internal teams do not think about which model handles what. They describe the job, and the gateway routes it. If one model degrades, we swap it. If a cheaper, faster model becomes available in-house, we route to it. We are not locked into one giant that can hold us hostage on price or capability. We own the routing layer. We own the workforce. The complexity is ours, and so is the leverage.
The people who argue for one large model are usually arguing for simplicity of management, not simplicity of economics. In India, economics wins. A model that costs ₹50 a call does not survive our market. A model that takes 3 seconds to respond does not survive our customers. The gateway lets us be complex where it matters and simple where it counts.
What this means
- Many small AI models vs one large is not a debate—it is a cost statement. The large model loses on unit economics in India every time.
- Specialization is the only path to ≤800ms voice-to-voice. No single giant model hits that floor while staying under ₹10 a call.
- The gateway is the moat, not the models. Whoever controls routing controls cost, latency, and quality.
- Content and voice are the same problem. The Video Engine and our voice agents share one philosophy: small, specialized, ruthlessly routed.
- If you are building for India, build for the price floor first. Everything else—quality, speed, scale—has to fit underneath ₹6–10 a call or it does not ship.