Applied AI / LLM Engineer building evaluated agent systems, RAG pipelines, and production AI tooling.
I build production-minded AI systems with Python, LangGraph, LangChain, OpenAI, Anthropic, Streamlit, SQLite, vector search, and API backends. My work focuses on retrieval reliability, agent memory, evaluation, observability, privacy boundaries, tool safety, and cost-aware execution.
I am a 2026 B.Tech graduate and AI Engineering Fellow at Maven (AI Makerspace), open to Applied AI / LLM Engineer roles with Indian and international AI teams.
- Built a live MCP agent-security demo with GitHub Pages, CI, policy evaluation, tool-risk scoring, redaction, and OpenRouter LLM explanations.
- Engineered DevMind with six security-aware tools, persistent sessions, runtime metrics, plugin support, CI, and 156 tests.
- Built ContextOps Agent with typed memory, plan persistence, context compression, privacy review, measurable evals, and observed CI success.
- Built OpenAI AutoData with challenger/solver/judge agents, persistent budget controls, fail-closed validation, auditable outputs, 13 regression tests, and CI.
- Built research-backed RAG systems with corrective retrieval, adaptive routing, citation grounding, precision/recall/F1, and RAG evaluation metrics.
- Completed Andrew Ng's five-course Deep Learning Specialization and continue studying production RAG, agent evaluation, and context engineering.
| Proof | Link |
|---|---|
| MCP Sentinel Lab live demo | prince2-ai.github.io/mcp-sentinel-lab |
| MCP Sentinel Lab repository | github.com/PRINCE2-AI/mcp-sentinel-lab |
| Latest MCP Sentinel CI run | GitHub Actions |
| Latest MCP Sentinel Pages deploy | GitHub Pages deploy |
| Project | What it does | Engineering signals |
|---|---|---|
| MCP Sentinel Lab | Runtime security gateway and evaluation bench for MCP/tool-using AI agents | Live demo, policy engine, risk scoring, redaction, OpenRouter explanations, CI, GitHub Pages |
| EvoCode Scientist | AlphaEvolve-inspired coding agent that evolves and evaluates candidate solutions | Evolution loop, benchmark scoring, sandboxed execution, lineage tracking, Streamlit/API surface |
| DevMind | Terminal-native AI coding agent built with Python, LangGraph, and Claude | Six built-in tools, persistent sessions, runtime metrics, plugins, cross-platform support, 156 tests, CI |
| ContextOps Agent | Context-engineering layer for long-horizon agents with typed memory, compression, and privacy review | Plan persistence, memory graph reconstruction, privacy firewall, token-savings metrics, API, dashboard, CI |
| Secure RepoPilot | Issue-to-PR coding agent with baseline verification, safety controls, and privacy auditing | Reproducer, minimal patching, command guardrails, leakage audit, SWE-style judge, API, dashboard, CI |
| TrustDI Agentic RAG | Agentic RAG system for trustworthy enterprise data integration and schema matching | Adaptive routing, evidence-backed decisions, OpenAI explanations, precision/recall/F1 evals, API, dashboard, CI |
| OpenAI AutoData | Generates hard research QA data through challenger, solver, and judge agents | Persistent budget guard, fail-closed validation, auditable outputs, 13 regression tests, CI |
- Long-horizon agents with explicit memory, context budgets, and quality gates
- RAG systems with retrieval evaluation, citation grounding, and failure recovery
- Tool-using agents with privacy boundaries, observability, and cost controls
- AI evaluation for faithfulness, reliability, safety, and regression testing
- Python APIs and deployable AI applications
- Corrective Agentic RAG Assistant: Adaptive CRAG assistant with query routing, corrective retrieval actions, hierarchical retrieval, web fallback, and RAG metrics.
- MemoryOS Agent: MemGPT-inspired long-term memory agent with OpenAI API support, SQLite memory, lifecycle controls, and Streamlit dashboard.
- Adaptive RAG CAG Project: adaptive RAG and cache-augmented generation demo for retrieval workflows.
- Local RAG with Ollama and ChromaDB: private PDF question answering with local inference and embeddings.
- AI Reddit Brand Monitor: local sentiment, topic, urgency, and feedback analysis for Reddit mentions.
Python LangGraph LangChain OpenAI API Anthropic API RAG Vector Search Streamlit FastAPI SQLite PyTorch Docker GitHub Actions RAGAS-style evals
- Build the smallest reliable system that proves the idea.
- Test failure paths, not only happy paths.
- Make cost, state, and model behavior visible.
- Keep claims aligned with reproducible code and results.
- Document limitations clearly instead of overstating benchmark performance.
I am open to Applied AI / LLM roles and focused open-source collaboration. Technical feedback is welcome.
Best repositories to pin for recruiter review: MCP Sentinel Lab, EvoCode Scientist, DevMind, ContextOps Agent, Secure RepoPilot, and TrustDI Agentic RAG.
