Semih Yavuz
Research Director
My background
Semih Yavuz is a Research Director at Salesforce AI Research, leading a team focused on improving the factuality, groundedness, and reasoning capabilities of large language models in knowledge-intensive applications. His work involves developing state-of-the-art embedding and re-ranker models for knowledge retrieval across diverse domains, including code, multi-modal, and multilingual contexts, while refining retrieval-augmented generation (RAG) by enhancing how LLMs consume and integrate knowledge in complex reasoning. His team is focused on pushing the boundaries of the research to develop accurate, scalable, and reliable AI systems and driving product impact with them in the CRM domain.
Semih's latest articles
Blog
Post-training methods (RLVR, On-policy distillation) are Episode-local Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewards (RLVR), a model tries a problem, a verifier checks…
Blog
SFR-VibeTrain: The Agent That Trains Agents
What if launching an RL training run felt less like operating a GPU cluster and more like talking to a sharp research engineer in Slack? Training AI models is still strangely artisanal, involving…
4 authors
Blog
Building Efficient RL Training for the Agentic Era
Introduction Reinforcement Learning from Human or AI Feedback (RLHF, RLAIF) has become the standard recipe for aligning large language models (LLMs). But as we push into the agentic era — where models call…
3 authors
Blog
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
VIBEPASS, a new benchmark, reveals a fundamental weakness in modern AI coding assistants: even with near-perfect scores on code generation tasks, frontier models falter when it comes to finding and fixing subtle bugs…
3 authors
Blog
TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation
The era of software engineering agents is underway. Benchmarks and real-world usage (e.g., tools like Cursor and Claude Code) illustrate that LLMs can be incredibly effective at writing code for real-world use-cases. Over…
3 authors