I am a Member of Technical Staff at Inception, working from our Bangalore office across the training and inference stack for Mercury, our family of diffusion LLMs. Previously I was a Research Scientist at Adobe Research, working on NLP and multi-modal machine learning.
I received my PhD from the Department of Computational and Data Sciences (CDS) at the Indian Institute of Science, Bangalore, advised by Prof. Partha Talukdar, where I worked on knowledge graph embeddings for question answering (thesis).
Prior to graduate school, I received my Bachelor of Engineering (BE Hons.) in Computer Science from Birla Institute of Technology and Science, Pilani (BITS Pilani) in 2015.
Research
My work comes down to two questions: how do you make language models faster without giving up quality, and how do you verify that what a model generates is actually supported by what you gave it.
On speed, I wrote Prompt Lookup Decoding, a speculative decoding method now integrated into Hugging Face Transformers and vLLM, and I currently work on inference for diffusion LLMs at Inception.
On grounding, I work on attribution — reading a model’s attention to trace generated text back to the exact source spans behind it. That research shipped in Adobe Acrobat’s AI Assistant and is now the foundation of TokenPath.
During my PhD I worked on knowledge graph embeddings for question answering. During my undergraduate studies I worked on humanoid robotics and computer vision.
A full publication list is on Google Scholar.
Projects
Everything 3D
Get custom 3D-printed figurines and personalized keepsakes made from your photos. Order a unique gift from the store or follow the latest creations on Instagram :)
Store: everything3dindia.com
Instagram: everything_3d_india
TokenPath.ai
TokenPath is an API for grounding AI-agent outputs by tracing generated tokens back to the exact source tokens behind them. It grew out of my research on attribution in document-grounded question answering, and provides token-level attribution signals for precise citations, provenance, hallucination gates, and evals.
Website: tokenpath.ai
Docs: docs.tokenpath.ai
Experiments
Knowledge Cutoff Benchmark
A simple benchmark to measure the actual knowledge cutoff of LLMs, rather than the date labs report. The results are surprising: OpenAI and Anthropic are the only major labs keeping their models fresh — most others, including the Chinese labs, lag by 12+ months.
Retrieval vs Recall
Agentic search benchmarks may not measure search. Running WideSearch with retrieval turned off, one frontier model scored slightly better without it — the tasks are built from facts old enough to answer from memory. Rebuilt on post-cutoff events, the ranking changed. Write-up published on the Inception blog.
The Seen and the Unseen — Transcripts
A searchable browser for Whisper-generated transcripts of Amit Varma’s podcast The Seen and the Unseen, covering 300+ episodes. Read any episode’s transcript in a clean, per-episode view.
Selected publications
Saxena A. Prompt Lookup Decoding.
Developed a method to speed-up LLM decoding, integrated in transformers and vLLM.
D.J. Bajpai, S. Agarwal, A. Saxena, K. Kulkarni, S. Mitra & M.K. Hanawal. “FlowCast: Trajectory Forecasting for Scalable Zero-Cost Speculative Flow Matching”. International Conference on Learning Representations (ICLR 2026).
S. Somasundaram, A. Phukan & A. Saxena. “PLD+: Accelerating LLM Inference by Leveraging Language Model Artifacts”. Findings of the Association for Computational Linguistics: NAACL 2025.
A. Phukan, S. Somasundaram, A. Saxena, K. Goswami & B.V. Srinivasan. “Peering into the mind of language models: An approach for attribution in contextual question answering”. Findings of the Association for Computational Linguistics: ACL 2024, 11481-11495. Shipped in Adobe Acrobat’s AI Assistant.
Saxena A., Kochsiek A. & Gemulla R. “Sequence-to-Sequence Knowledge Graph Completion and Question Answering”. Accepted to the 2022 Annual Conference of the Association for Computational Linguistics (ACL 2022).
Saxena A., Chakrabarti S. & Talukdar P. “Question Answering Over Temporal Knowledge Graphs”. Accepted to the 2021 Annual Meeting of the Association for Computational Linguistics (ACL 2021).
Saxena A., Tripathi A. & Talukdar P. “Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base Embeddings”. Accepted to the 2020 Annual Conference of the Association for Computational Linguistics (ACL 2020).
Work Experience
Inception, Bangalore
Member of Technical Staff — September 2025 - Present
Working across the training and inference stack for the Mercury series of diffusion LLMs.
- Core algorithmic work on application-specific decoding performance, including a 2x speedup on Mercury Edit
- Led post-training of Mercury 2 for an enterprise query-rewrite workload — quality parity with their production Gemini baseline, 2.5x faster
- Tool-calling data and evaluation for the Mercury 2 and Mercury 2.5 releases
- Work with search and answer-engine customers on agentic retrieval pipelines
Adobe Research, Bangalore
Research Scientist — June 2022 - August 2025
Document Experiences Lab — research on LLMs for document understanding, generation and grounding, taken from idea to shipped product.
- Shipped attribution into Acrobat’s AI Assistant, giving users sentence-level grounding for answers over long documents
- Contributor to the first Adobe Firefly image model
- 15+ papers at ACL, EMNLP, NAACL, ICCV, EACL and INLG; 12+ patents filed
- Mentored interns and junior researchers, several of whom led first-author papers at ACL, NAACL and EMNLP
Google, Hyderabad
Software Engineer, Tools and Infrastructure — May 2016 - Mar 2017
PayPal, Chennai
Software Engineer 1 — July 2015 - November 2015
Teaching
Guest lecturer on LLM decoding in Introduction to Natural Language Processing (DS 207) at the Indian Institute of Science, invited two years running (2025 and 2026). Course by Prof. Danish Pruthi.
CV
Contact
apoorv [at] inceptionlabs.ai
apoorvumang [at] gmail.com