MMManoj Mali

PRINCIPAL AI ENGINEERING · PUNE, INDIA

AI engineering.
At production scale.

I’m Manoj. I design agentic AI systems and lead their path to production—connecting retrieval, tools and evaluation with the architecture needed for reliable, efficient products.

Currently at Gruve · Since April 2025Previously Symbl.ai, Chalo & Crowdfire

13+ years building
systems behind products.

Agentic systemsRAG & retrievalEvaluation & LLMOpsTechnical leadership

01 / SELECTED WORK

AI architecture.
Production outcomes.

Selected work in real-time AI, platform scale and data-intensive products.

01

CONVERSATION INTELLIGENCE

Scale the platform.
Control the cost.

Distributed systemsCloud infrastructure

At Symbl.ai, I worked on the infrastructure behind a conversation intelligence platform processing millions of interactions.

60 → 5,000+concurrent sessions

I scaled concurrent workloads through distributed architecture optimization and reduced infrastructure costs by 50%. Engineering practices introduced during this work reduced production defects by 20%.

Explore the work

Engineering contribution

Platform architecture and infrastructure optimization, alongside engineering leadership on RAG systems and real-time AI assistants.

Why it matters

Capacity and cost determine how far a real-time AI product can grow. This work connected platform improvements to both operating efficiency and customer-facing capability.

Career timeline
02

REAL-TIME AI

Put context inside
the conversation.

PythonRAGChromaDBWebSocket

AI assistance is most useful when it arrives while a decision is still being made.

My portfolio includes a RAG-based assistant that provides contextual suggestions during live sales calls, using vector search, knowledge-base integration and WebSocket communication.

Live conversationContext retrievalRelevant guidance
Explore the work

Implementation focus

Python, FastAPI, ChromaDB and Redis support retrieval and real-time delivery. A related call analytics project combines streaming transcription, NLP and automated scoring with Kafka, LangChain and OpenAI.

Product capability

Context-aware guidance during calls, plus structured analysis and performance metrics for sales teams.

03

MOBILITY PLATFORMS

Build for a city
in motion.

JavaMicroservicesCI/CD

At Chalo, I worked on backend services for public transport and ticket booking.

1M+GPS records processed daily

I transformed monolithic services into microservices and built a ticket-booking platform serving 40,000+ monthly active users.

Explore the work

Engineering contribution

Service decomposition, ticketing backend development and deployment automation. CI/CD pipelines reduced manual deployment intervention by 80%.

Why it matters

Transport products depend on continuously changing location data. The backend needs to support that flow while keeping product delivery manageable for the engineering team.

Career timeline

02 / EXPERIENCE

Principal experience.
AI engineering focus.

Previously Principal Software Engineer at Symbl.ai. Today, I’m a Staff Software Engineer at Gruve, with a focus on principal-level AI engineering opportunities.

View LinkedIn
22 April 2025–Present

Gruve

Staff Software Engineer

2020–2025

Symbl.ai

Principal Software Engineer

2017–2020

Chalo

Software Engineer III

2015–2017

Crowdfire

Senior Software Developer

Article recommendation and social analytics systems
2013–2015

Tata Consultancy Services

System Engineer

03 / TOOLS AND THINKING

From agent design
to production ownership.

Agentic systems

Multi-agent orchestration, planning and handoffs, tool/function calling, MCP, memory and context management, prompt engineering, LangGraph, LangChain, LlamaIndex.

Retrieval and grounding

Document ingestion, chunking, embeddings, hybrid search, reranking and grounded generation. RAG pipelines with ChromaDB and pgvector.

Evaluation and experimentation

Offline and online evaluation, golden datasets, regression gates, model benchmarks, LLM-as-judge, agent task success and human feedback loops.

LLMOps and observability

Prompt and model versioning, experiment tracking, tracing and performance monitoring. LangSmith, Langfuse, OpenTelemetry, Prometheus and Grafana.

Model adaptation and serving

PyTorch, Hugging Face, SFT, LoRA/QLoRA and PEFT; vLLM and OpenAI, Anthropic and Gemini APIs. Model selection across quality, latency and cost.

Reliability and voice AI

Guardrails, red-teaming, safe tool execution, approval gates, retries, timeouts and durable workflows. Streaming ASR, turn detection and barge-in.

Architecture and leadership

Technical strategy, roadmaps, ADRs, design reviews, mentoring and cross-functional alignment. Translate business metrics into engineering priorities.

Production engineering

Async Python, FastAPI, Java, Node.js, SQL, Kafka, Redis and PostgreSQL. AWS, GCP, Kubernetes, Docker, Terraform, CI/CD and GitOps.

LET’S TALK ENGINEERING

Building an AI product
that needs to scale?

I’m interested in Principal AI Engineer roles at startups, with ownership of AI architecture, production delivery and the technical direction of the team.

manoj@manojmali.comDownload resume