Services
What I build
RAG Architecture (current focus)
Multi-stage retrieval
Design and implement a multi-stage retrieval architecture combining lexical and semantic retrieval:
- BM25 lexical retrieval and bi-encoder semantic retrieval
- Hybrid Search and Reciprocal Rank Fusion (RRF)
- Cross-encoder reranking to improve retrieval precision
- Query expansion using synonym maps and query rewriting
- Multi-hop retrieval workflows for complex information needs
- Full orchestration via LangChain and LangGraph, with explicit state management and controlled execution flow
- Evaluation using Recall@k, MRR, and nDCG@k
- Validation, evaluation, and guardrail mechanisms for reliability and predictability
Broader AI Engineering
- Design and implement LLM-based applications and agentic AI systems
- Deploy and integrate local open-source models (Hugging Face, PyTorch, vLLM), including inference optimization and resource-aware deployment
- Parameter-efficient fine-tuning (PEFT/LoRA), including adapter-based model specialization
- Agentic workflows with tool calling, MCP, memory, structured outputs, and multi-agent orchestration
- Software architecture principles applied to AI: modularity, maintainability, observability, scalability, clear separation between deterministic and non-deterministic components
- Evaluation of models, architectures, frameworks, and deployment approaches against practical constraints: quality, latency, memory, scalability, complexity
Also available
I can also teach your team — because my knowledge goes deeper than just calling an API.
Available for project engagements, architecture reviews, or ongoing advisory work — get in touch to discuss what fits.