Skip to content
← All opportunities

Backend Engineer - Vision

JobBengaluruvia Ashby

Posted 18 Aug

Apply on Sarvam AI →

Opens the company's own listing — Aspirova never mirrors applications.

Sign in to save

Skills for this role

Skills we detected for this role.

  • Python
  • Docker
  • Kubernetes
  • PostgreSQL
  • Redis
  • Computer Vision
  • Generative AI
  • Apache Airflow
  • REST APIs
  • SQL

Opportunity brief

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Team

Sarvam's research teams build our own vision-language models for OCR and structured extraction. This team builds everything around them — the serving harness that turns a 3B or 30B in-house model into a production document intelligence platform.

The bet is specific: with the right harness — routing, decomposition, retries, verification, ensembling, layout awareness, confidence calibration — a small sovereign model should match or beat what teams today get from frontier hosted models like Gemini Flash, at a fraction of the cost and fully within India. Closing that gap is an engineering problem, and it is this team's problem.

We run against the full messiness of Indian documents at population scale: PAN and Aadhaar, bank statements, GST filings, insurance and medical reports, 60-page

contracts, legal filings and RFPs — across languages, scan quality, and layouts that were never designed to be machine-read.

Stack: Go, Python, Temporal, REST, Kubernetes, PostgreSQL, Redis, object storage, OpenTelemetry-based observability.

About the Role

You will build and own services inside the document intelligence harness — the layer that sits between our vision models and the enterprises consuming them. A single 100- page document turning into clean structured JSON involves ingestion, page-level fan out, model inference, post-processing, validation, and assembly, all of which has to be durable, observable, and fast. You will own pieces of that path end to end.

This is a build role early in your career, but not a scoped-ticket role. You will be given real surface area, reviewed closely, and expected to grow into owning services outright.

What You'll Do

Build and maintain RESTAPIs for document submission, job status, and result retrieval — both synchronous and long-running async flows

Implement Temporal workflows and activities for multi-stage document pipelines: split, pre-process, infer, post-process, validate, assemble

Write the pre- and post-processing that decides real accuracy: page segmentation, de-skewing, layout handling, schema validation, output normalisation Instrument everything — latency and cost per stage, per page, per model — so that regressions are visible before customers find them

Build internal tooling: replay harnesses, evaluation runners, golden-set regression suites for extraction accuracy

Debug production issues across the stack: stuck workflows, GPU queue backpressure, malformed inputs, partial failures

Work directly with the models team to turn model behaviour into harness behaviour

What We're Looking For

  • 1–2 years building backend services that have actually run in production Strong fundamentals in Go or Python — you can write clean, tested, concurrent code and reason about what it does under load
  • Solid grasp of HTTP and RESTAPI design, and of async/background job processing Comfort with PostgreSQL and Redis; you understand transactions, indexing, and where state should live
  • Working familiarity with Docker; exposure to Kubernetes
  • You debug by reading logs, traces and code — not by guessing
  • Genuine interest in AI systems engineering; you want to go deep on making models work in production, not on calling APIs

Bonus Points

Exposure to Temporal or another durable workflow engine (Airflow, Cadence, Step Functions)

Any hands-on work with OCR, computer vision, or LLM/VLM inference Familiarity with GPU serving stacks — vLLM, TensorRT-LLM, Triton, SGLang Open-source contributions or side projects we can read

Note

We are looking for people who can own the outcomes described here, not people who match every line of this specification. If this problem excites you and you believe you can do this work, we want to hear from you.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers ofAI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier ofAI in India, Sarvam is the place to be.

More like this

Related opportunities