Rui Meng's current work focuses on designing state-of-the-art embedding models and benchmarks for text, code, and image search. Rui also builds retrieval-augmented generation (RAG) and re-ranking models. His research interests lie in how both humans and models memorize, represent and understand data in general.
The SFR-Embedding-Mistral marks a significant advancement in text-embedding models, building upon the solid foundations of E5-mistral-7b-instruct and Mistral-7B-v0.1.
TLDR We trained a series of 7B LLMs named XGen-7B with standard dense attention on up to 8K sequence length for up to 1.5T tokens. We also fine tune the models on public-domain…