An Exam for Active Observers
Researchers have introduced ActiveVision, a new benchmark designed to test whether multimodal large language models can perform active observation like humans.
- Current MLLM benchmarks lack the ability to measure active, iterative visual perception.
- ActiveVision introduces 17 tasks across 3 categories to simulate human-like gaze redirection.
- The benchmark focuses on the 'closed loop' nature of vision rather than static image analysis.
Current vision-language benchmarks often rely on static snapshots, which fail to capture the dynamic nature of human vision. Human sight is a continuous loop where gaze is redirected based on evolving hypotheses, a process known as active observation.
ActiveVision addresses this gap by introducing 17 distinct tasks across three categories. These tasks require models to engage in repeated visual perception to solve complex problems, moving beyond simple image-to-text matching.
By measuring how well models can navigate and interact with visual information through iterative gaze, this benchmark provides a more rigorous test for the next generation of multimodal large language models (MLLMs).
Provides a more rigorous testing framework for multimodal model training and evaluation.
Offers a new research direction in cognitive-inspired AI evaluation.
Moves AI closer to how humans actually perceive and interact with the world.
- MLLM
- Multimodal Large Language Model, an AI capable of processing multiple types of data like text and images.
- Active Observation
- The process of continuously redirecting gaze based on intermediate hypotheses to better understand a scene.
Doctors Develop Guiding Principles for Future of AI in Healthcare - UVA Health
Preparing Communities for AI Risks Facing Older Adults at #MACoCon - Conduit Street Blog
AI in 2026: Smarter Models, Harder Questions - USC Viterbi School of Engineering
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
These are the most urgent AI risks, according to 272 experts - MIT Sloan
The Business Students Who Want to Use AI for Good - USC Today
USC business students are using AI to drive positive change in their communities.
BusinessX relaunches a rebuilt Android app after year-long effort
X has released a rebuilt version of its Android app after a year-long development effort.
BusinessOpenAI is scared of open-weight models. Should the US be?
OpenAI's concerns about open-weight LLMs from China have sparked a US policy debate.
Survey shows bipartisan support for federal AI safety regulations - Washington Examiner
A recent survey indicates bipartisan support for federal regulations on AI safety in the US.
Transportation looks to AI to accelerate its modernization initiatives - Nextgov/FCW
The US transportation sector is exploring the use of artificial intelligence to accelerate its modernization initiatives.

China’s AI models have Trump’s AI world at war with itself
Current and former Trump advisors publicly criticized leading US AI companies, arguing they are failing to counter the threat posed by China's AI models.