Srijan Bansal
Applied Scientist, AI Research
My background
Srijan Bansal is an Applied Scientist specializing in LLM reasoning, long-context understanding, and automatic evaluation. He has contributed to the development of XGen-Sales and is actively involved in SWE Agent research and automated testing applications. He earned his Master’s from Carnegie Mellon University.
My expertise
Generative Models, RAG, Agent
When I'm not working
I enjoy outdoor activities like tennis, swimming, surfing, and biking. I also love going on drives and cooking.
Srijan's latest articles
Blog
Post-training methods (RLVR, On-policy distillation) are Episode-local Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewards (RLVR), a model tries a problem, a verifier checks…
Blog
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
VIBEPASS, a new benchmark, reveals a fundamental weakness in modern AI coding assistants: even with near-perfect scores on code generation tasks, frontier models falter when it comes to finding and fixing subtle bugs…
3 authors
Blog
TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation
The era of software engineering agents is underway. Benchmarks and real-world usage (e.g., tools like Cursor and Claude Code) illustrate that LLMs can be incredibly effective at writing code for real-world use-cases. Over…
3 authors
Blog
Software Engineering Agent via Self-Abstraction from Grounded Experience
Large language model (LLM)-based software engineering (SWE-) agents have recently demonstrated remarkable progress on realistic software engineering tasks such as code review, bug fixing, and repository-level reasoning. Most SWE-agents start from a fresh…
9 authors
Blog
Does Context Matter? Introducing ContextualJudgeBench for RAG and Summarization Evaluation
AI is rapidly transforming industries, helping businesses enhance customer experiences, improve efficiency, and make smarter decisions. But an essential question arises: How can we ensure that AI is creating accurate and grounded answers?…
5 authors