Hi, I’m Manish.

I lead GenAI engineering at UsefulBI, building multi-agent systems, their evals and their memory.

I also build open-source tools to understand how agents work, write about evals, and once topped a CUDA kernel leaderboard.

Open to workSenior or Lead AI Engineer roles · Bangalore, remote or relocation

Agents & tools Evals & memory Notes & posts A few side quests

A little about me.

I’ve been in this field since the classical ML era, back when “AI” meant feature engineering, LayoutLM and hand-tuned pipelines. These days it means agents, and the evals and memory that make them worth trusting.

Before industry I did speech and vision research at IISc Bangalore. I like building small tools to find out how things actually work, and sharing them in the open.

Find me on LinkedIn
Manish Sharma
Manish, in Bangalore.

Projects

Some start with a problem at work.
Some with “how does that actually work?”

Manish Code planning a todo list, reading calc.py and fixing it with str_replace.Manish Code refusing to delete a .env file because the permission policy blocks it.

Manish Code

Manish Code is a Claude Code-style coding agent I built from scratch in plain Python. One readable agent loop drives file, shell, edit, planning and skill tools over any OpenAI-compatible model.

The interesting work is in keeping it safe: a permission engine that blocks destructive commands and bypass tricks, a kernel-level OS sandbox with no network, and automatic context compaction so long sessions don’t fall over. Every file is small enough to read.

OpenMemoryUI: a chat beside session, episodic, semantic and contextual memory panels.OpenMemoryUI title card: transparent agentic memory.

OpenMemoryUI

What is an agent actually remembering? OpenMemoryUI turns agent memory into a glass box: session, episodic, semantic and contextual memory update in front of you as you chat.

The useful part is being able to inspect every step. Each turn runs through a six-stage pipeline, from receive to write, with retrieval scores, prompt inspection, provenance and edit history. It connects to live models and MCP servers, or runs a zero-setup demo.

Try the demo
Tab Bouncer after checking 12 tabs for a Lisbon trip: 5 rabbit holes found, 9 tabs to close.

Tab Bouncer

One line about what you’re doing, and a door policy for every open tab. Tab Bouncer is a Chrome extension that shows the tabs you don’t need the way out.

Under the hood there’s no agent loop and no scraping. It makes a single call to Jev, TypeSafe AI’s decision model, with your task and each tab’s title and address, and gets back a score and a category for up to 120 tabs in about a second, for a hundredth of a cent.

OpenMCP UI: say what you need, MCP handles the handoff.OpenMCP UI flow: you ask, the client chooses a tool, the server runs it.

OpenMCP UI

The Model Context Protocol is easy to name and surprisingly hard to picture. OpenMCP UI is a visual playground: ask in plain English and watch the client pick a tool, prepare its arguments and hand the call to the server.

Every step of the exchange is traced, from the handshake and tools/list to the final tools/call. A no-code workshop generates your own MCP servers, and it talks to OpenAI, Anthropic, OpenRouter and Gemini.

Watch a tool call
Video RAG showing timestamp windows and five retrieved video frames of planets.Video RAG answer with visual and textual evidence and timestamps.

Video RAG

Can you cite a moment in a video the way you’d cite a page? Video RAG indexes a YouTube video in Qdrant, then answers questions with the frames and timestamps it used as evidence.

The answer view shows the frames it retrieved next to the answer, so you can check its reasoning instead of taking it on trust. There’s a short Loom walkthrough, too.

Experience

Building AI that has to work on Monday morning.

UsefulBI

Leading GenAI engineering

As Lead AI Engineer, since August 2025, I lead enterprise GenAI platform work: multi-agent orchestration, no-code workflow builders, evaluation systems, quality intelligence and AWS/EKS production deployment.

  • Gen AI Studio. The CSR multi-agent orchestration layer on AWS and Strands: modular agent workflows, task routing and reusable orchestration patterns.
  • No-code agent builder. A drag-and-drop canvas for composing configurable agent pipelines, reused across CSR, PLPS and DSUR flows.
  • Evals and regression gating. The eval harness for the Decision Tree Agent on Strands and Bedrock: golden datasets, tool-call trajectory graders, span-level tracing of tokens, cost and latency, a human-calibrated LLM-as-a-Judge, pass@k and pass^k checks, and a gate that blocks releases that underperform the baseline.
  • Online evals. Reviewer decisions become production labels, giving a continuous AI-human agreement metric with no extra instrumentation.
  • Agent memory. Role-scoped short-term memory in DynamoDB, long-term episodic memory of past quality events in a vector database, and an audit trail for every override.
  • IQ Quality Intelligence. Similar-event retrieval over historical quality events with multimodal evidence, with end-to-end ownership of the User, FAM and QE workflows, backend and AWS/EKS deployment.
  • AI platform and data. Architected the SUNY course discovery platform and the GQMD unified data repository across Smartsheet Data Vault, GPLM and Databricks pipelines.
  • Leadership. Led 4 engineers, ran 10+ AI interviews, organised a company-wide pharma AI hackathon and curated 50+ AI/ML learning resources.
4
engineers led
10+
AI interviews
50+
AI/ML resources
3
agent flows: CSR, PLPS, DSUR
More about UsefulBI
Parspec

Datasheets at scale

As ML Engineer II, from January 2024 to July 2025, I built production ML and LLM systems for construction-product intelligence: datasheet retrieval, attribute extraction, annotation automation and multimodal RAG evaluation.

  • Datasheet intelligence. Model-number to datasheet mapping across 5M documents, 80% faster with DynamoDB caching and Gemini 2.0 Flash.
  • Fine-tuning. Llama 3.3 70B Instruct on H100s (Modal) for structured attribute extraction.
  • Extraction quality. Family-name recall from 89% to 96% with Gemini 2.5 Flash; header and column detection from 4 to 8 columns at 97% accuracy.
  • LLM-as-a-Judge. GPT-4o and Gemini in parallel to triage datapoints for human annotation, cutting manual annotation to about 74%.
  • Multimodal RAG. BGE text and CLIP image embeddings in FAISS: 0.94 MRR in retrieval and 0.96 accuracy in generation.
  • Infrastructure. Migrated the AI workloads to AWS EKS, with API Gateway and isolated staging and production.
80%
latency reduction
96%
recall
0.94
retrieval MRR
More about Parspec
Docsumo

ML Scientist, Dec 2022 to Dec 2023. Document KV and table extraction with LayoutLM, BROS and YOLO in production, integrated into 10+ client APIs supporting $80K-$100K MRR. Built Chat-AI with LangChain and Pinecone, and cut annotation from a full day to about two hours with GPT-4 extraction.

IISc

Research Assistant, Speech and Vision, 2021 to 2022. OCR and speech-recognition pipelines at SpireLabs, Indian Institute of Science: Hindi text and audio datasets, PyTesseract and EasyOCR with word-level accuracy, and Librosa MelSpectrogram audio work.

B.Tech in Electronics and Communication, NMIT Bangalore, GPA 8.98. The full story is in my resume.

Download resume

Skills

The stack behind the work.

Agent engineering

  • Agents
  • Multi-agent collaboration
  • Agent orchestration
  • Task routing
  • Tool calling
  • Reusable agent nodes
  • Workflow composition
  • No-code agent builder
  • Agent deployment

Agent evaluation

  • Agent tracing
  • Engineering harness
  • Evaluation harness
  • Experiment tracking
  • Comparative testing
  • LLM-as-judge
  • Tool-call evaluation
  • Compliance scoring
  • Response fidelity
  • Accuracy checks
  • Task completion
  • Trajectory grading
  • Regression gating
  • Online evals

RAG and retrieval

  • RAG
  • Multi-agent RAG
  • Multimodal RAG
  • Vector search
  • Embeddings
  • Similar-event retrieval
  • Qdrant
  • FAISS
  • Pinecone

Models and frameworks

  • GPT-4o
  • Gemini
  • Llama 3.3 70B
  • PyTorch
  • Hugging Face
  • LangChain
  • LangGraph
  • LlamaIndex
  • Strands Agent

Document AI and NLP

  • OCR
  • LayoutLM
  • BROS
  • YOLO
  • Table extraction
  • KV extraction
  • SpaCy
  • SciSpacy
  • MedXN
  • RxNORM
  • Librosa

Cloud and deployment

  • AWS
  • EKS
  • ECS
  • Bedrock
  • API Gateway
  • Docker
  • Kubernetes
  • Netlify
  • Production deployments
  • Staging/prod partitions

Data pipelines

  • Databricks
  • Data pipelines
  • Smartsheet Data Vault
  • GPLM architecture
  • DynamoDB
  • S3
  • GCS
  • MySQL
  • Neo4j
  • Hasura

Backend and product

  • Python
  • FastAPI
  • Flask
  • TypeScript
  • Next.js
  • Model Context Protocol
  • API design
  • Backend architecture
  • Workflow builders
  • Annotation workflows
  • User/FAM/QE flows
  • Linux

Builder tools

  • Cursor
  • Ollama
  • OpenRouter
  • Modal Labs
  • Google Colab
  • Kaggle
  • GitHub
  • Tableau
  • Amplitude
  • Lovable

In my notes

A few things I’ve written about agents.

  1. Demos show pass@k. Users live in pass^k.

    Why a 90% agent can still fail most five-run checks.

  2. An open-source handbook on AI evals, from the basics to interview questions

    Free for everyone.

  3. Agent memory’s hard problem isn’t storage. It’s contradiction.

  4. Open-sourcing a mini Claude Code harness in plain Python

On the leaderboard

Side quests with a scoreboard.

GPU MODE runs a competition where you write and optimise CUDA kernels for AI workloads. I hadn’t touched GPU programming in about three years.

Instead of hand-tuning, I ran a tight research, benchmark and iterate loop: an autoresearch-style harness that kept testing changes for two hours on a $20 plan. It reached rank #1. Sometimes the best way to relearn something is to build the loop that learns with you.

Visit GPU MODE

Loop engineering

Research, benchmark, iterate, repeat. The same idea behind my eval harnesses, pointed at a kernel instead of an agent.

Read the write-up

Got an agent that
needs to work? Let’s talk.

I’m open to contract projects and full-time roles in applied AI: agentic systems, evals, RAG and document intelligence. A messy prototype, a pipeline that needs gating, or an idea that needs a first working version are all welcome starting points.

01

Build a multi-agent system that ships.

02

Evaluate and gate an agent before release.

03

Rescue a RAG or document AI pipeline.

manish.tinkering@gmail.com LinkedIn · Manish Sharma Twitter / X · @lucifer_x007