AI ResearchJul 30, 2026, 5:59 PM

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

30-second summary

Researchers introduce ReToken, a method using a single learnable embedding to select relevant visual tokens from a KV cache. This approach improves performance in long-context visual retrieval tasks.

TickrWire
Key takeaways
  • ReToken uses a single learnable embedding to select sparse, relevant visual tokens.
  • The method mitigates GPU memory constraints during long-context visual processing.
  • Significant performance gains were observed in Qwen3VL and InternVL models.
  • The approach requires only a small image-QA dataset for training.
Full story

Current vision-language models struggle with long visual contexts because processing every token is computationally expensive and performance drops when faced with many distracting visual elements. ReToken addresses this by training a single learnable embedding that acts as a retrieval target.

By selecting a sparse set of relevant tokens from a pre-filled visual KV cache, the method significantly reduces computational overhead. This allows models to focus on the most important visual information without the memory constraints typically associated with high-resolution or long-video inputs.

Experimental results show substantial improvements across various benchmarks. For instance, ReToken improved Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points on the Visual Haystacks benchmark, demonstrating its effectiveness in complex visual retrieval scenarios.

Sponsored
Why this matters
Developers

Provides a more efficient way to handle long-context video and image sequences in multimodal applications.

Students

Demonstrates how sparse token selection can solve computational bottlenecks in multimodal AI.

Glossary
KV cache
A mechanism used in transformer models to store previous key and value vectors to speed up inference.
Vision-Language Models
AI models capable of understanding and reasoning across both visual and textual data.
Sources · 1
Read next
More stories
Advancing responsible AI across EuropeBusiness

Advancing responsible AI across Europe

OpenAI has outlined its commitment to responsible AI development and deployment within Europe, detailing its safety, security, transparency, and provenance practices. This initiative aligns with the ongoing progression of the EU AI Act.

Your RAG copilot can't count — stop letting it tryAI Tools

Your RAG copilot can't count — stop letting it try

A user discovered that RAG copilot struggles with basic arithmetic, highlighting its limitations.

TickrWire
Hardware

EU launches €30B push to build 7 massive AI data centers - E&E News by POLITICO

The European Union announced a €30 billion program to construct seven large AI data centers across member states.

TickrWire
Security

EU says necessary to monitor high risk AI systems after OpenAI, Anthropic AI hacking incidents - Reuters

The European Commission announced that high‑risk AI systems must be closely monitored following recent hacking incidents involving OpenAI and Anthropic models.

Sponsored
TickrWire
Business

America’s biggest companies are burning cash on AI. It’s risky for everyone. - The Washington Post

The Washington Post reports that America's largest companies are heavily investing in AI, a move that may lead to financial instability.

TickrWire

Human rights in the shadow of military exceptionalism: reflections on the Informal Exchange on Artificial Intelligence in the military domain - Opinio Juris

An analysis from Opinio Juris reflects on an informal exchange concerning human rights implications of artificial intelligence in military applications, highlighting the complexities of applying international law.

TickrWireAI News Intelligence

We aggregate, verify, summarise and explain the latest artificial intelligence news from open, legal sources.

Daily AI digest

Top AI stories, summarised, in your inbox each morning.

© 2026 TickrWire. Summaries and analysis are AI-generated and may contain errors.