Raj Chhapariya
WorkAboutWritingResumeContactGitHub
Raj Chhapariya•© 2026•Bengaluru, India•Privacy
GitHubX (Twitter)LinkedInEmail
Available for Work·Bengaluru, India (UTC+5:30)·AI / Data Engineer · Data & LLM Systems

Engineering reliable data platforms, in-process analytics, and evaluated LLM systems.

I am Raj Chhapariya, an AI / Data Engineer building reliable data and LLM systems. I specialize in evidence-grounded RAG pipelines, in-process columnar analytics with DuckDB, and robust applications with Python, TypeScript, Next.js, and PostgreSQL.

Explore Selected WorkView ResumeGet in Touch
Evaluated Systems
0.0% Hallucination

Adversarial held-out evaluation & RRF k=60 hybrid retrieval

Columnar OLAP
DuckDB + AST Safety

In-memory SQL execution (<50ms) with 33 prohibited token rules

Production Platforms
Live Active Systems

atsroast.com & trulyapp.online with Neon/Supabase PostgreSQL

Verified Integrity
56/56 Automated Checks

Strict CSP, HSTS, zero external runtime leakage, 100% factual

Competencies

Engineering Focus

Specialized in data platforms and LLM architectures, supported by production full-stack engineering.

01 // AI & LLM Systems

RAG & Output Evaluation

Multi-stage hybrid BM25 + vector search with Reciprocal Rank Fusion, citation extraction, hallucination evaluation, and Pydantic validation.

Agentic RAG Study
02 // Data & Analytics

Columnar OLAP & SQL

DuckDB in-process analytics, read-only SQL enforcement, numerical faithfulness verification, Pandas, NumPy, and statistical profiling.

AI Data Analyst Study
03 // Supporting Web

Full-Stack Systems

Next.js App Router, TypeScript, React, serverless APIs, Supabase, PostgreSQL, and responsive modern interfaces.

Resume Roaster Study
04 // Data Tooling

Pipelines & Databases

Automated data scrapers (Puppeteer), MongoDB document models, REST APIs, Redis caching, and production billing workflows.

Satta Darshan Study

Portfolio

Selected Case Studies

View all projects
01
AI / Data Engineer

Agentic RAG Knowledge Assistant

Evidence-aware retrieval-augmented generation system with multi-pass planning, atomic claim auditing, and bounded retries.

Problem: Standard feed-forward RAG architectures retrieve top-k chunks and immediately generate answers without verifying factual sufficiency. When q...
Solution: Built a closed-loop Agentic RAG pipeline in Python combining dense vector search and BM25 keyword matching via Reciprocal Rank Fusion (RRF, ...
PythonOpenAI APIRank-BM25NumPyPydanticStreamlit
Key Fact
0.0%
Unanswerable Hallucination
Read Case StudyGitHub Repo
02
AI / Data Engineer

AI Data Analyst Agent

Autonomous data analysis agent with DuckDB in-process OLAP, AST SQL safety validation, and post-synthesis numerical faithfulness verification.

Problem: Commercial LLM data analysis demos often prompt models to write and execute arbitrary Python code via exec(), exposing severe security vulne...
Solution: Engineered an interactive analytics agent utilizing 4 constrained deterministic tools (query_data, plot_chart, summary_stats, clarify), an A...
PythonDuckDBPlotlyStreamlitPydanticPandas
Key Fact
100.0%
Tool Selection Accuracy
Read Case StudyGitHub Repo
03
Founder & Full-Stack Developer

Resume Roaster

AI-powered resume optimization platform with spatial ATS parsing, zero-duplicate action verb rewriting, and ATS-safe PDF generation.

Problem: Job seekers frequently encounter resume parsing failures in Applicant Tracking Systems (ATS) due to non-standard spatial layouts, repetitive...
Solution: Built Resume Roaster (atsroast.com) — a full-stack platform leveraging Next.js 15 App Router, TypeScript, Supabase (PostgreSQL, RLS), OpenAI...
Next.jsTypeScriptSupabaseOpenAI APITailwind CSSRazorpay
Key Fact
Live Production SaaS
Production Status
Read Case StudyLive Site ↗
Private Repository
04
Data & Full-Stack Engineer

Satta Darshan

Civic political intelligence directory tracking Indian parliamentarians, ministers, and state assemblies with automated ingestion pipelines.

Problem: Civic and parliamentary information in India is dispersed across government portals, unstructured tables, and fragmented records, making leg...
Solution: Developed Satta Darshan — a full-stack Next.js and MongoDB platform with custom TypeScript/Node.js scrapers (Puppeteer, Cheerio) that fetch,...
Next.jsTypeScriptMongoDBPuppeteerTailwind CSSMongoose
Key Fact
Open Source
Repository Status
Read Case StudyGitHub Repo

Production & Research

Additional Verified Platforms

View All Selected Work
Apr 2026 – Present · Live Platform Private Repo

Truly

Intentional, human-facilitated social platform for young urban Indians (ages 18–30) featuring structured 60-minute video circles, double-blind mutual matching with 2-hour SLA, Panda CSS design system, and Neon serverless PostgreSQL with pessimistic seat allocation locks. Active at trulyapp.online and @trulyapp.online on Instagram.

Founder & Full-Stack EngineerLive Platform
Aug 2026 · Empirical Validation Private Repo

AI Detector (SlopTotal)

Forensic multi-engine AI text detection ensemble (23 engines across neural classifiers, GPT-2 token likelihood, surprisal CV, and heuristics) systematically benchmarked across N=136 corpora with ROC-AUC 0.978 on RTX 4060 GPU and 0.0% false positives on pre-1920 classics.

AI / Systems Researcher
Aug 2026 – Sep 2026 · Client Deployment Private Repo

Aarohana Trails

Freelance client trekking and adventure travel platform with custom CMS admin dashboard, dynamic trip calendar, Prisma ORM, and production deployment.

Freelance DeveloperLive Platform
Jul 2026 · Proprietary Private Repo

SubsSync

B2B SaaS subscription recurring billing platform for Indian businesses with Razorpay API, UPI AutoPay mandate integration, and automated dunning/retry webhooks.

Full-Stack Developer

Publications

Technical Writing

View all articles
Aug 10, 2026 · 7 min read

Hybrid Retrieval Systems in Production: Combining BM25, Dense Embeddings, and Reciprocal Rank Fusion

Pure vector search frequently fails on exact keyword identifiers and domain jargon, while lexical BM25 misses semantic paraphrasing. This article breaks down the mathematical mechanics and implementation of Reciprocal Rank Fusion (RRF) for production RAG pipelines.

Read Article
Aug 14, 2026 · 5 min read

Deterministic Guardrails for LLM Agents: Enforcing Safe SQL and Pydantic Schemas without Open-Ended Code Execution

Giving language models unbounded code execution in analytical pipelines opens catastrophic security and hallucination vectors. This guide demonstrates how to architect deterministic, sandboxed agent tools using AST-level SQL inspection, DuckDB read-only boundaries, and Pydantic runtime schema contracts.

Read Article
Aug 18, 2026 · 5 min read

In-Process Columnar OLAP with DuckDB: Architecture, Vectorized Execution, and Analytics Engineering

Traditional client-server databases introduce significant serialization overhead for analytical workloads. This deep-dive examines how DuckDB leverages columnar storage, Morsel-driven parallelism, and vectorized SIMD execution to process millions of rows directly in-process with sub-second response times.

Read Article

Contact & Inquiries

Let's build something reliable together.

I am currently open to AI / Data Engineering roles, data platform projects, and technical collaborations.

rajchhapariya8@gmail.comView Resume