Opportunity brief
What You'll Own:
- Inference infrastructure: Stand up and operate the inference engine, deploying vLLM as the primary solution and evaluating TGI where model coverage requires it, across GPU capacity from India-region partners and burst/global providers.
- API gateway: Build and maintain the OpenAI-compatible API surface, including /v1/chat/completions, /v1/embeddings, and /v1/models, so existing OpenAI SDK code can run against MIRA with a one-line endpoint change.
- Smart routing: Build MIRA's core differentiator, including the complexity classifier that scores incoming requests and routes them across Tier 1–3 models, the post-generation quality validation step, and the escalation path to frontier fallback providers such as Anthropic/OpenAI. Implement the logging required to tune routing thresholds using real traffic.
- Model catalogue: Deploy and manage the v1 model catalogue, including Llama, Qwen, Deepseek, Mistral, Mixtral, nomic-embed-text, and BGE-M3 embeddings, and evaluate Indic models such as Sarvam-1 for the roadmap.
- Platform integration: Integrate MIRA with auth.thq.digital SSO and the unified THQ credit wallet, including per-key usage metering for input/output tokens, model, tier, and escalation flag, which will serve as the billing source of truth.
- Observability & unit economics: Implement per-request logging, latency and error dashboards, and per-customer GPU cost attribution from day one. This data will support sustainable pricing decisions.
- Pilot delivery: Support benchmark data collection for the MeiTY 90-day pilot, including latency P50/P95, INR cost-per-million-tokens compared with AWS Bedrock/OpenRouter/Together AI, and OpenAI conformance testing. Contribute to the public Day-90 report.
- Compliance awareness: Build with India's DPDP Act requirements in mind, including data residency by default and explicit consent flagging for any frontier fallback call that leaves Indian infrastructure.
What We're Looking For:
- 1–3 years of experience building and shipping production ML/backend systems. The experience requirement is flexible for candidates with strong, demonstrable inference/LLM-serving expertise.
- Strong Python skills and experience with FastAPI or similar frameworks for building API services.
- Solid understanding of async I/O and request handling at scale.
- Hands-on experience with vLLM, TGI, or comparable LLM-serving frameworks.
- Practical understanding of PagedAttention-style batching, quantization trade-offs, and GPU memory and throughput tuning.
- Experience building or fine-tuning lightweight classification models, or the ability to build a rules-based heuristic and evolve it over time.
- Understanding of designing and validating tiered routing and escalation systems.
- Experience with Docker and basic Kubernetes or equivalent orchestration.
- Experience working with cloud GPU providers such as AWS, GCP, CoreWeave, Nebius, E2E, or Neysa.
- Experience with Prometheus, Grafana, or equivalent observability tools.
- Understanding of what needs to be logged for accurate cost attribution and billing, beyond basic uptime monitoring.
- Ability to read a PRD, identify gaps, and make sound engineering decisions independently in an environment with ambiguity around pilot timelines, GPU capacity, and pricing.
Nice to Have:
- Exposure to OpenAI API conformance testing or building drop-in replacement API surfaces.
- Familiarity with India-specific requirements, including the DPDP Act, GST-compliant billing flows through platforms such as Razorpay/Cashfree, and data residency requirements.
- Interest in or exposure to Indic-language models such as Sarvam and IndicBERT.
- Prior experience in an early-stage or founder-led environment where you were the first engineer on a product line.
Why This Role:
- Opportunity to work as the founding engineer on a product with a live government pilot track through MeiTY/IndiaAI Mission and a defined commercial thesis.
- Build against an existing MIRA PRD and a defined architecture rather than starting from a blank page.
- Direct access to the founder, ownership of technical decisions, and an opportunity to grow into THQ.Digital's broader AI/ML platform work as MIRA develops.