Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
Researchers introduce Group-Calibrated On-Policy Distillation ( practices) to fix errors in long-context reasoning during model training.
- practices addresses the mismatch between local token guidance and global task success.
- practices uses task-specific verifiers to provide graded rewards for response-level completion.
- The method improves performance in evidence-aggregation tasks across long input ranges.
Current on-policy distillation methods often fail when handling long-context tasks because they rely on token-level guidance from a teacher model. This approach can lead the student model to prioritize locally plausible text Base while ignoring global constraints or evidence spread across a large input iter.
- On-policy distillation
- A training method where a student model learns from its own generated responses using guidance from a larger teacher model.
- The process of training a model based on the distribution of its own outputs.
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Finetuning Strategies for Querying Sounds by Vocal Imitation
Interpretable AI predicts a 2026 summer dry anomaly in central China
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
Algeria adopts roadmap for sovereign artificial intelligence - Muslim Network TV
Algeria has officially adopted a national roadmap to develop sovereign artificial intelligence capabilities, aiming to reduce reliance on foreign AI systems.
Israeli AI-based start-up Dondy acquired by UK holding company Circeus - The Jerusalem Post
UK-based Circeus has acquired Dondy, an Israeli AI startup, marking another strategic move in the global AI consolidation trend.
BusinessStripe didn’t really buy OpenRouter because of the ‘singularity’
Stripe has acquired OpenRouter, an AI model routing startup, to enhance its AI capabilities for payments processing and fraud detection.
Turkcell Advances 6G Technologies and Artificial Intelligence R&D - The Fast Mode
Turkcell has announced new advancements in 6G technology and AI research, positioning itself as a leader in next-generation wireless and intelligent systems.
SecurityI Built an AI Code Reviewer. Then OWASP Broke It.
An experiment shows popular AI coding assistants struggle with OWASP security rules, failing 70% of tests in a code review scenario.
BusinessOpenAI seeks to one-up Anthropic with new customer privacy protections
OpenAI introduces stricter privacy controls for enterprise customers, aiming to surpass Anthropic's existing protections.