Mohammed Usmani

Mohammed Usmani · AI EngineerI build agentic AI that runs in production.

AI Engineer at Garage, Bengaluru. Multi-agent orchestration, retrieval, realtime voice, and the multi-tenant backends behind them.

Ask This Portfolio

Ask anything about my work — answered in real time by an agent that searches my actual resume and projects, with sources.

Try

experience

AI & ML Engineer

Garage (Gravitichain Technology Group Pvt. Ltd.) · Bengaluru, India

Feb 2026 — Present
  • Cut the sales copilot's worst-case prompt from ~96K to ~2K tokens per turn with a seven-layer prompt stack.
  • Made spreadsheet bugs unrepresentable: the LLM emits a typed IR, and a deterministic compiler writes the workbook.
  • Turned a single-user agent runtime into a multi-tenant service without forking it: 265 tools, 57 integrations.
  • Shipped a crash-recoverable meeting note-taker that cuts 20–40% of input tokens before summarising.
+ more systems → full resume

selected work

The model never writes a cell address01 · INPUTCA request02 · MODELLLM03 · TYPED IRinputs · rows · scenariosNO CELL-ADDRESS SURFACE04 · COMPILERdeterministic05 · OUTPUT.xlsx✕ LLM writes formulascross-sheet ref bugs: possible

Luca

CA assistant where the LLM emits a spreadsheet IR, never cell formulas.

● live · IR → compiler, no formulas

Seven layers instead of stuffing the context~96K~2K

NetworkChains

Sales copilot cutting worst-case context from ~96K to ~2K tokens per turn.

● live · ~96K → ~2K tokens/turn

The LLM points at cells. Arithmetic decides the link.01 · LINKS TO EARNStatement linepresented figureNotebreakdownTB accounttrial balance02 · RESOLVERdeterministic03 · TIE-OUT GATEarithmeticNO EDGE BYPASSES ITPASSEDLLM → LLM_RECONCILEDFAILEDnot reconciledNEEDS_REVIEWnot a false FAILEDON MISSLLM rescueextracts figuresPROVENANCELLM_CANDIDATEnot yet confirmedCELL REMAPreal Cellfigure → cell

FS Preparation Agent

Financial statements where the LLM points at cells and arithmetic decides every link.

1,828 tests

Only the Controller can end a task01 · VOICEuser taskPLANPlannersub-goals onlyON FAILEDre-planmax 1SUB-GOALSsuccess_criteriamax_steps 5 each02 · PERCEIVEscreen stateXML + screenshot03 · ~19 ACTIONSExecutorgets screenshot04 · DEVICEUI actionAccessibilityServiceSETTLE250msno half-drawn frameCHECK · FLASH-LITEVerifierXML only05 · CONTROLLERowns terminationONLY EXIT06 · ANSWERTTS answerre-read from screen

EyesAI

An Android agent for blind users where only the Controller, never the model, declares success.

only the Controller ends a task

all 23 systems → /work

stack · 69 technologies

shipped in production work

Languages5

LLMs & agents12

Retrieval & RAG10

Voice & realtime9

Backend10

Data8

ML & vision3

Cloud, infra & frontend12

recognition

Top 17 of 5000+AI For Good Hackathon2025
3rd PrizeGNEC International Hackathon2024
3rd PrizeCodeathon DSA Hackathon2024
all certificates →

faq

Who is Mohammed Usmani?

Mohammed Usmani is an AI Engineer based in Bengaluru, India, currently at Garage (Gravitichain Technology Group). He builds production agentic AI systems — multi-agent orchestration, retrieval pipelines, realtime voice, and the multi-tenant backends behind them — and has shipped 15+ of them across sales, accounting and workspace products.

What AI systems has he actually built?

Four stand out. A financial-statement engine whose Cross-Reference Graph links every statement line to its note and ledger account, confirmed by arithmetic tie-out rather than model judgement. NetworkChains, an AI sales platform combining an agentic copilot, a realtime call assistant, a relationship graph and conversational image editing. Luca, a ~199,000-line AI accounting platform for Indian Chartered Accountants running on five LLM providers. And EyesAI, an autonomous Android agent that operates a phone end-to-end for blind and low-vision users.

What is his experience with RAG and retrieval engineering?

He built a seven-layer prompt assembly over six purpose-scoped Qdrant collections that cut worst-case input from roughly 96,000 to 2,000 tokens per turn. The retrieval goes well beyond nearest-neighbour search: contextual retrieval that prepends an LLM-written blurb before embedding, HyDE expansion for terse queries, reciprocal rank fusion with a 14-day recency half-life, hybrid dense and lexical search, and LLM reranking with diversification. He has also worked with pgvector, FAISS and Pinecone.

Has he built realtime voice AI?

Yes. He built a live in-meeting copilot that captures the host microphone and remote participants as two separate streams, downsamples 48kHz to 16kHz PCM16 in 250-millisecond batches, and gives each stream its own Deepgram socket — so speaker attribution comes from the transport rather than a diarization model. He also built a voice-to-voice restaurant ordering system on Pipecat and Gemini Live over a full-duplex WebSocket, and telephony voice agents over Telnyx SIP.

What is the Cross-Reference Graph?

It is the core data structure of the financial-statement engine: 9 node kinds and 6 edge kinds modelling how a set of financial statements actually hangs together. Links are made by deterministic resolvers and then confirmed by arithmetic tie-out, and every edge records how it was created and whether the numbers reconcile. Where an LLM assists, it only extracts and points at cells — the pass/fail verdict comes from the same deterministic arithmetic, so no reconciliation passes without a verified tie-out.

Which programming languages and tools does he use?

Primarily Python with FastAPI and TypeScript with Node, Express and NestJS. Day to day that means LangGraph, the Model Context Protocol, OpenAI, Anthropic Claude and Gemini, Deepgram and Voxtral for speech, Qdrant and pgvector for retrieval, Celery and BullMQ for background work, PostgreSQL, MongoDB and Redis for storage, and Docker with GitHub Actions deploying to GCP, AWS and DigitalOcean.

Is he available for hire, and where is he based?

Yes — he is open to AI/ML and backend engineering roles at product-focused startups. He is based in Bengaluru, India, and can be reached at mohammedusmani2005@gmail.com or on LinkedIn.

Building agents that have to work?

I'm looking for AI / backend engineering roles at product-focused teams.

mohammedusmani2005@gmail.com