ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Researchers introduce ReToken, a method using a single learnable embedding to select relevant visual tokens from a KV cache. This approach improves performance in long-context visual retrieval tasks.
- ReToken uses a single learnable embedding to select sparse, relevant visual tokens.
- The method mitigates GPU memory constraints during long-context visual processing.
- Significant performance gains were observed in Qwen3VL and InternVL models.
- The approach requires only a small image-QA dataset for training.
Current vision-language models struggle with long visual contexts because processing every token is computationally expensive and performance drops when faced with many distracting visual elements. ReToken addresses this by training a single learnable embedding that acts as a retrieval target.
By selecting a sparse set of relevant tokens from a pre-filled visual KV cache, the method significantly reduces computational overhead. This allows models to focus on the most important visual information without the memory constraints typically associated with high-resolution or long-video inputs.
Experimental results show substantial improvements across various benchmarks. For instance, ReToken improved Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points on the Visual Haystacks benchmark, demonstrating its effectiveness in complex visual retrieval scenarios.
Provides a more efficient way to handle long-context video and image sequences in multimodal applications.
Demonstrates how sparse token selection can solve computational bottlenecks in multimodal AI.
- KV cache
- A mechanism used in transformer models to store previous key and value vectors to speed up inference.
- Vision-Language Models
- AI models capable of understanding and reasoning across both visual and textual data.
How Artificial Intelligence Discovered A New Way To Detect Patients At Risk Of Cardiac Death Using Simple EKGs - Forbes
First AI-driven telescope goes stargazing - Northwestern Now News
China's MiniMax releases H3 video model - Reuters
AI ResearchHow a Baseten Engineer Traced 7 Years of Attention Mechanism Evolution -- From GPT-2 to Kimi K3, in Runable PyTorch
Can one screening strategy find many cancers? Artificial Intelligence is bringing the idea closer - EurekAlert!
BusinessAdvancing responsible AI across Europe
OpenAI has outlined its commitment to responsible AI development and deployment within Europe, detailing its safety, security, transparency, and provenance practices. This initiative aligns with the ongoing progression of the EU AI Act.
AI ToolsYour RAG copilot can't count — stop letting it try
A user discovered that RAG copilot struggles with basic arithmetic, highlighting its limitations.
EU launches €30B push to build 7 massive AI data centers - E&E News by POLITICO
The European Union announced a €30 billion program to construct seven large AI data centers across member states.
EU says necessary to monitor high risk AI systems after OpenAI, Anthropic AI hacking incidents - Reuters
The European Commission announced that high‑risk AI systems must be closely monitored following recent hacking incidents involving OpenAI and Anthropic models.
America’s biggest companies are burning cash on AI. It’s risky for everyone. - The Washington Post
The Washington Post reports that America's largest companies are heavily investing in AI, a move that may lead to financial instability.
Human rights in the shadow of military exceptionalism: reflections on the Informal Exchange on Artificial Intelligence in the military domain - Opinio Juris
An analysis from Opinio Juris reflects on an informal exchange concerning human rights implications of artificial intelligence in military applications, highlighting the complexities of applying international law.