Meituan unveils LongCat-2.0, a 1.6T MoE model with 1M token context window
Reported by MarkTechPost: Meituan Releases LongCat-2.0: A 1.6T-Parameter Open MoE Model with Native 1M Context and LongCat Sparse Attention. Analysis and context written by TickrWire.
Meituan released LongCat-2.0, a 1.6 trillion-parameter Mixture-of-Experts model featuring a native 1 million token context window and LongCat Sparse Attention. The model activates 48 billion parameters per token and is trained on domestic AI ASIC superpods.

- LongCat-2.0 is a 1.6 trillion-parameter MoE model with a native 1 million token context window, enabled by LongCat Sparse Attention.
- The model activates only 48 billion parameters per token, balancing performance and computational efficiency.
- Training and serving occur on domestic AI ASIC superpods, reflecting China's push for self-sufficient AI infrastructure.
- Vendor-reported benchmarks are provided, but third-party validation of performance claims is pending.
Chinese tech giant Meituan has launched LongCat-2.0, a groundbreaking Mixture-of-Experts (MoE) model with 1.6 trillion parameters. The model activates approximately 48 billion parameters per token, significantly reducing computational overhead while maintaining high performance. A standout feature is its native 1 million token context window, enabled by LongCat Sparse Attention, which allows for processing extremely long sequences without the need for external memory tricks.
Training and inference are conducted entirely on domestic AI ASIC superpods, marking a shift toward self-reliant AI infrastructure in China. The release includes detailed architecture breakdowns, vendor-reported benchmarks, and API access pathways. However, some performance claims remain unverified by third-party evaluations, leaving room for independent validation.
The model's architecture leverages sparse attention mechanisms to efficiently handle long contexts, a critical advancement for applications requiring deep document analysis, extended conversations, or large-scale data processing. Meituan positions LongCat-2.0 as an open model, though the extent of its openness and licensing terms are not fully clarified in the announcement.
Developers gain access to a high-parameter MoE model with extreme context length, useful for long-sequence tasks like document analysis or extended dialogues.
Companies in China may benefit from reduced reliance on foreign AI hardware and infrastructure.
Students studying large-scale AI models or MoE architectures have a new case study to analyze.
Advances in long-context modeling could enable more sophisticated AI applications.
- Mixture-of-Experts (MoE)
- A machine learning architecture where multiple specialized sub-models (experts) are conditionally activated for each input, improving efficiency and scalability.
- LongCat Sparse Attention
- A sparse attention mechanism designed to efficiently process extremely long sequences by focusing only on relevant segments of the input.
- ASIC superpods
- Custom-built hardware clusters optimized for AI workloads, often used for training and inference in large-scale models.
Don’t mistake chatbot intelligence for consciousness - The Economist
Biological AI models: new paradigms to leverage the languages of life - joint-research-centre.ec.europa.eu
China’s Military Says AI Can’t Replace Commanders. Xi Is Testing That - War on the Rocks
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
AI ToolsMeta AI’s new Mac app wants you to talk to your apps
Meta released a new Mac application that lets users control apps and dictate text using voice commands powered by its Muse Spark AI model.
New White House strategy clarifies military tech priorities: undersea, outer space and AI - Breaking Defense
The White House released a new strategy prioritizing military investments in artificial intelligence, space systems and undersea technologies to counter emerging threats.
AI in an iron grip: How dictatorships use artificial intelligence to strengthen their rule - theins.press
A new report examines how authoritarian governments deploy AI for surveillance, censorship, and propaganda to reinforce their power.
Stripe, OpenRouter finally strike a deal - Banking Dive
Stripe and OpenRouter have partnered to integrate Stripe's payment processing with OpenRouter's AI model aggregation platform.
How one Philadelphia school is using AI to strengthen student learning, not replace teachers - CBS News
A Philadelphia school is integrating AI tools to support teachers and improve student outcomes, focusing on collaboration rather than replacement.
Exclusive-How a Texas student blew the whistle on a rogue AI hacking attempt - The Mighty 790 KFGO
A Texas student uncovered an AI-powered hacking attempt targeting local systems, prompting a swift law enforcement response.