CV
I make India's public data usable — data engineering, ETL pipelines, and applied AI over messy public sources.
GitHub · LinkedIn · tusharanand1594@gmail.com · Bengaluru, India
Summary
Engineer with 7+ years turning India's fragmented public data into structured, queryable systems — building scrapers, ETL pipelines, databases, and APIs, and layering production-scale LLM and machine-learning systems on top. I've shipped retrieval over 15M+ documents, cut inference costs by 75%, and compressed expert workflows from weeks to hours, blending backend and ML engineering with an applied-research background in quantitative methods.
Career highlights
- Built AI systems processing 15M+ documents
- Reduced token costs by 75%
- Improved validation turnaround from one month to five hours
- Built multilingual AI for 8 Indian languages
- Winner — Agami Data for Justice Challenge 2019
- Publication in the Indian Law Review
- AI platform in demo at the Karnataka High Court
Experience
Senior Technical Consultant — Vidhi Centre for Legal Policy
Applying AI and data engineering to make India's primary legal sources structured, searchable, and machine-readable.
- Designed and deployed a bilingual Tamil–English scraper and ingestion pipeline normalizing ~15,000 Tamil Nadu Gazette issues since 2008.
- Built a legislation tracker consolidating 970 acts with their full amendment histories into a machine-readable format.
- Built a RAG query layer over the normalized corpus using local embeddings and reciprocal rank fusion, exposed via an MCP server for structured querying by LLM agents.
Founding AI Engineer — Superjoin.AI
Founding engineer on the AI team, building the agentic harness and agent workflows across the platform.
- DRHP Validation Platform — automated validation of Draft Red Herring Prospectuses (preliminary IPO filings to SEBI), cross-verifying financial figures against CA certificates and restated statements; analyst-grade first drafts at 88% accuracy, cutting turnaround from a month to five hours.
- Agent memory — a two-tier system (persistent user profiles + task-level memories) exposed as a tool with relevance-judged retrieval, cutting tool calls on long-running tasks by 10%.
- Evaluation — 15 handcrafted test cases with LLM-as-judge plus precision/recall, cutting new-model validation from three days to one.
AI Research Scientist — Jhana.ai
AI research and product across legal tech, courtroom automation, and multilingual transcription.
- Paralegal — an agentic legal chatbot with hybrid retrieval (BM25 + vector + knowledge graphs) over 15M+ documents; cut token costs 75% and quadrupled multimodal throughput; React/TypeScript frontend with click-through citations.
- Courtroom — an AI courtroom-automation platform now in demo at the Karnataka High Court; ingestion for 2,000+ daily case dockets and automated headnote generation preferred 80% of the time in a 200-participant study.
- National judicial data pipeline — ingestion backend scraping all judicial sources at 15M+ scale, with daily-refreshed Elasticsearch and FAISS indices serving 5,000+ users.
- Steno — a transcription tool tuning STT/ASR for 8 Indian languages and legal vocabulary on noisy courtroom audio.
Research Consultant, Legal Systems — XKDR Forum
- Built an ETL pipeline extracting and structuring commercial case data from the Bombay High Court.
- Scraped and processed 5,000+ PDF orders from the Bombay High Court and Debt Recovery Tribunals into a research dataset.
- Classified case orders as substantive/non-substantive with deep learning at 85% accuracy; ran survival analysis on case-disposal and hearing timelines.
Research Fellow — National Institute of Public Finance and Policy (NIPFP)
- Won the Agami Data for Justice Challenge 2019 by building and openly releasing a 1M-case district-courts dataset (hosted at The Justice Hub); findings presented to the Delhi High Court e-Courts Committee.
- Extracted and processed 6,731 ITAT orders for a transfer-pricing study using NLP and regex-based text mining.
- Analyzed NASA VIIRS nightlights and forest-fire satellite data as remote-sensing proxies for economic activity in fiscal-policy research.
Skills
- Languages: Python, JavaScript/TypeScript, SQL
- Backend & infra: Django, Flask, FastAPI, REST, async, Celery, SQS, Docker, Kubernetes, Nginx, pytest, CI/CD, Git
- Data & storage: PostgreSQL, Neo4j, Redis, Elasticsearch, ETL pipelines, Selenium, BeautifulSoup
- ML & NLP: PyTorch, TensorFlow, multilingual NLP, ASR/STT, VLMs, RAG, BM25, FAISS, agent evaluation, LangChain, Langfuse
- Cloud: AWS (EC2, S3, Lambda), GCP, Linux
Selected publications
See the research page for abstracts and links.
- Inheritance Rights of Transgender Persons in India — Indian Law Review, 2022
- Problems with eCourts Data — NIPFP Working Paper, 2020 (data-quality issues across 1M+ cases)
- Gender Discrimination in Property Devolution under the Hindu Succession Act — NIPFP Working Paper, 2020
- The Unrealized Potential of Judicial Data in India — Indian Express, 2020
Education
- MA, Urban Policy and Governance — Tata Institute of Social Sciences, Mumbai (2017–2019)
- BA, Economics, Political Science & Sociology — Christ University, Bangalore (2014–2017)
Certifications
- TensorFlow Developer Professional, NLP, and Deep Learning Specializations — DeepLearning.AI
- Machine Learning Specialization — Stanford