Agentic RAG Knowledge Assistant
Evidence-aware retrieval-augmented generation system with multi-pass planning, atomic claim auditing, and bounded retries.
I am Raj Chhapariya, an AI / Data Engineer building reliable data and LLM systems. I specialize in evidence-grounded RAG pipelines, in-process columnar analytics with DuckDB, and robust applications with Python, TypeScript, Next.js, and PostgreSQL.
Adversarial held-out evaluation & RRF k=60 hybrid retrieval
In-memory SQL execution (<50ms) with 33 prohibited token rules
atsroast.com & trulyapp.online with Neon/Supabase PostgreSQL
Strict CSP, HSTS, zero external runtime leakage, 100% factual
Competencies
Specialized in data platforms and LLM architectures, supported by production full-stack engineering.
Multi-stage hybrid BM25 + vector search with Reciprocal Rank Fusion, citation extraction, hallucination evaluation, and Pydantic validation.
DuckDB in-process analytics, read-only SQL enforcement, numerical faithfulness verification, Pandas, NumPy, and statistical profiling.
Next.js App Router, TypeScript, React, serverless APIs, Supabase, PostgreSQL, and responsive modern interfaces.
Automated data scrapers (Puppeteer), MongoDB document models, REST APIs, Redis caching, and production billing workflows.
Portfolio
Evidence-aware retrieval-augmented generation system with multi-pass planning, atomic claim auditing, and bounded retries.
Autonomous data analysis agent with DuckDB in-process OLAP, AST SQL safety validation, and post-synthesis numerical faithfulness verification.
AI-powered resume optimization platform with spatial ATS parsing, zero-duplicate action verb rewriting, and ATS-safe PDF generation.
Civic political intelligence directory tracking Indian parliamentarians, ministers, and state assemblies with automated ingestion pipelines.
Production & Research
Intentional, human-facilitated social platform for young urban Indians (ages 18–30) featuring structured 60-minute video circles, double-blind mutual matching with 2-hour SLA, Panda CSS design system, and Neon serverless PostgreSQL with pessimistic seat allocation locks. Active at trulyapp.online and @trulyapp.online on Instagram.
Forensic multi-engine AI text detection ensemble (23 engines across neural classifiers, GPT-2 token likelihood, surprisal CV, and heuristics) systematically benchmarked across N=136 corpora with ROC-AUC 0.978 on RTX 4060 GPU and 0.0% false positives on pre-1920 classics.
Freelance client trekking and adventure travel platform with custom CMS admin dashboard, dynamic trip calendar, Prisma ORM, and production deployment.
B2B SaaS subscription recurring billing platform for Indian businesses with Razorpay API, UPI AutoPay mandate integration, and automated dunning/retry webhooks.
Publications
Pure vector search frequently fails on exact keyword identifiers and domain jargon, while lexical BM25 misses semantic paraphrasing. This article breaks down the mathematical mechanics and implementation of Reciprocal Rank Fusion (RRF) for production RAG pipelines.
Giving language models unbounded code execution in analytical pipelines opens catastrophic security and hallucination vectors. This guide demonstrates how to architect deterministic, sandboxed agent tools using AST-level SQL inspection, DuckDB read-only boundaries, and Pydantic runtime schema contracts.
Traditional client-server databases introduce significant serialization overhead for analytical workloads. This deep-dive examines how DuckDB leverages columnar storage, Morsel-driven parallelism, and vectorized SIMD execution to process millions of rows directly in-process with sub-second response times.
Contact & Inquiries
I am currently open to AI / Data Engineering roles, data platform projects, and technical collaborations.