Mohammed Usmani · AI EngineerI build agentic AI that runs in production.
AI Engineer at Garage, Bengaluru. Multi-agent orchestration, retrieval, realtime voice, and the multi-tenant backends behind them.
Ask anything about my work — answered in real time by an agent that searches my actual resume and projects, with sources.
experience
AI & ML Engineer
Garage (Gravitichain Technology Group Pvt. Ltd.) · Bengaluru, India
- Cut the sales copilot's worst-case prompt from ~96K to ~2K tokens per turn with a seven-layer prompt stack.
- Made spreadsheet bugs unrepresentable: the LLM emits a typed IR, and a deterministic compiler writes the workbook.
- Turned a single-user agent runtime into a multi-tenant service without forking it: 265 tools, 57 integrations.
- Shipped a crash-recoverable meeting note-taker that cuts 20–40% of input tokens before summarising.
selected work
Luca
CA assistant where the LLM emits a spreadsheet IR, never cell formulas.
● live · IR → compiler, no formulas
NetworkChains
Sales copilot cutting worst-case context from ~96K to ~2K tokens per turn.
● live · ~96K → ~2K tokens/turn
FS Preparation Agent
Financial statements where the LLM points at cells and arithmetic decides every link.
1,828 tests
EyesAI
An Android agent for blind users where only the Controller, never the model, declares success.
only the Controller ends a task
stack · 69 technologies
Languages5
- Pythonshipped in FS Preparation Agent, Agent Orchestration Layer, NetworkChains
- TypeScriptshipped in Luca, NetworkChains, Garage, Agent Orchestration Layer
- Kotlinshipped in EyesAI
- JavaScriptshipped in CityFix
- SQL
LLMs & agents12
- Agentic tool-callingshipped in NetworkChains, Agent Orchestration Layer, EyesAI
- LangGraphshipped in NetworkChains
- LangChain
- MCPshipped in Agent Orchestration Layer
- OpenAIshipped in Garage, Luca
- Anthropic Claudeshipped in Garage, Luca
- Geminishipped in Garage, JobTrack, Luca, EyesAI
- Vertex AIshipped in EyesAI
- Azure OpenAIshipped in FS Preparation Agent, Luca
- Groqshipped in EyesAI
- Perplexityshipped in Luca
- Ollama
Retrieval & RAG10
- RAGshipped in Luca, NetworkChains
- Hybrid searchshipped in NetworkChains, Saffron & Smoke
- RRFshipped in Saffron & Smoke
- HyDEshipped in NetworkChains
- LLM rerankshipped in NetworkChains
- Cohere Rerankshipped in FS Preparation Agent
- Qdrantshipped in Agent Orchestration Layer, NetworkChains
- pgvectorshipped in FS Preparation Agent, Luca, Saffron & Smoke
- FAISSshipped in EyesAI
- Pinecone
Voice & realtime9
- Deepgramshipped in Garage, Luca, NetworkChains
- LiveKitshipped in Garage, NetworkChains
- Pipecatshipped in Saffron & Smoke
- Gemini Liveshipped in EyesAI, Saffron & Smoke
- Voxtral
- ElevenLabsshipped in Luca
- WebSocketsshipped in NetworkChains, Saffron & Smoke, Agent Orchestration Layer, EyesAI
- SSEshipped in Agent Orchestration Layer
- Silero VADshipped in Saffron & Smoke
Backend10
- FastAPIshipped in EyesAI, FS Preparation Agent, Agent Orchestration Layer, NetworkChains, Saffron & Smoke
- Node.jsshipped in Garage, NetworkChains
- Expressshipped in Luca, NetworkChains, Garage
- NestJSshipped in Soulmate
- Celeryshipped in Agent Orchestration Layer
- BullMQshipped in Garage, NetworkChains
- Socket.ioshipped in Soulmate
- SQLAlchemyshipped in FS Preparation Agent, Agent Orchestration Layer, Saffron & Smoke
- GraphQL
- REST
Data8
- PostgreSQLshipped in FS Preparation Agent, Luca, Saffron & Smoke, Agent Orchestration Layer
- MongoDBshipped in Garage, NetworkChains
- Redisshipped in Garage, NetworkChains, Saffron & Smoke, Agent Orchestration Layer
- Drizzleshipped in Luca, Soulmate
- Prismashipped in CityFix
- Supabaseshipped in GameWeb
- Firebaseshipped in CityFix, JobTrack, Soulmate
- Neo4j
ML & vision3
- TFLiteshipped in EyesAI
- ML Kitshipped in EyesAI
- Azure Document Intelligenceshipped in FS Preparation Agent
Cloud, infra & frontend12
- GCP
- AWS
- Railwayshipped in CityFix, FS Preparation Agent
- DigitalOcean
- Dockershipped in FS Preparation Agent, Agent Orchestration Layer
- GitHub Actions
- Nginx
- Reactshipped in Luca, Saffron & Smoke
- Next.jsshipped in FS Preparation Agent, Saffron & Smoke
- SolidJSshipped in SolidBoard
- Linux
- TailwindCSSshipped in GameWeb, Saffron & Smoke, Soulmate
recognition
faq
Who is Mohammed Usmani?
Mohammed Usmani is an AI Engineer based in Bengaluru, India, currently at Garage (Gravitichain Technology Group). He builds production agentic AI systems — multi-agent orchestration, retrieval pipelines, realtime voice, and the multi-tenant backends behind them — and has shipped 15+ of them across sales, accounting and workspace products.
What AI systems has he actually built?
Four stand out. A financial-statement engine whose Cross-Reference Graph links every statement line to its note and ledger account, confirmed by arithmetic tie-out rather than model judgement. NetworkChains, an AI sales platform combining an agentic copilot, a realtime call assistant, a relationship graph and conversational image editing. Luca, a ~199,000-line AI accounting platform for Indian Chartered Accountants running on five LLM providers. And EyesAI, an autonomous Android agent that operates a phone end-to-end for blind and low-vision users.
What is his experience with RAG and retrieval engineering?
He built a seven-layer prompt assembly over six purpose-scoped Qdrant collections that cut worst-case input from roughly 96,000 to 2,000 tokens per turn. The retrieval goes well beyond nearest-neighbour search: contextual retrieval that prepends an LLM-written blurb before embedding, HyDE expansion for terse queries, reciprocal rank fusion with a 14-day recency half-life, hybrid dense and lexical search, and LLM reranking with diversification. He has also worked with pgvector, FAISS and Pinecone.
Has he built realtime voice AI?
Yes. He built a live in-meeting copilot that captures the host microphone and remote participants as two separate streams, downsamples 48kHz to 16kHz PCM16 in 250-millisecond batches, and gives each stream its own Deepgram socket — so speaker attribution comes from the transport rather than a diarization model. He also built a voice-to-voice restaurant ordering system on Pipecat and Gemini Live over a full-duplex WebSocket, and telephony voice agents over Telnyx SIP.
What is the Cross-Reference Graph?
It is the core data structure of the financial-statement engine: 9 node kinds and 6 edge kinds modelling how a set of financial statements actually hangs together. Links are made by deterministic resolvers and then confirmed by arithmetic tie-out, and every edge records how it was created and whether the numbers reconcile. Where an LLM assists, it only extracts and points at cells — the pass/fail verdict comes from the same deterministic arithmetic, so no reconciliation passes without a verified tie-out.
Which programming languages and tools does he use?
Primarily Python with FastAPI and TypeScript with Node, Express and NestJS. Day to day that means LangGraph, the Model Context Protocol, OpenAI, Anthropic Claude and Gemini, Deepgram and Voxtral for speech, Qdrant and pgvector for retrieval, Celery and BullMQ for background work, PostgreSQL, MongoDB and Redis for storage, and Docker with GitHub Actions deploying to GCP, AWS and DigitalOcean.
Is he available for hire, and where is he based?
Yes — he is open to AI/ML and backend engineering roles at product-focused startups. He is based in Bengaluru, India, and can be reached at mohammedusmani2005@gmail.com or on LinkedIn.
Building agents that have to work?
I'm looking for AI / backend engineering roles at product-focused teams.
mohammedusmani2005@gmail.com