AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts LLM with 16B total parameters but only 2.8B active per token. The model was trained on Instinct MI300X and MI325X GPUs and includes full training weights, data mixtures, and inference code.

- Instella-MoE-16B-A3B is a fully open Mixture-of-Experts LLM with 16B total parameters but only 2.8B active per token, reducing computational load.
- The model was trained on AMD's Instinct MI300X and MI325X GPUs and includes full training weights, data mixtures, and inference code.
- AMD leverages Gated MLA and FarSkip-Collective techniques to enhance efficiency and performance in the MoE architecture.
- This release highlights AMD's commitment to open AI models and its growing presence in the AI hardware and software market.
AMD has introduced Instella-MoE-16B-A3B, a fully open Mixture-of-Experts (MoE) language model featuring 16 billion total parameters but activating just 2.8 billion per token. This design leverages Gated MLA and FarSkip-Collective techniques to optimize efficiency while maintaining performance. The model was trained from scratch on AMD's Instinct MI300X and MI325X GPUs, positioning it as a high-performance option for developers seeking open alternatives in the MoE space.
The release includes comprehensive resources: AMD has published the model weights from every training stage, along with the data mixtures used, configuration files, and inference code. This transparency aims to foster community adoption and further research, aligning with the growing demand for open and reproducible AI models. The move also underscores AMD's push into the AI hardware and software ecosystem, competing directly with established players in the LLM market.
Provides a fully open, efficient MoE model with transparent training resources, enabling customization and research.
Offers a competitive open alternative to proprietary MoE models, potentially reducing costs and dependency on closed solutions.
Signals AMD's strategic expansion into AI hardware and software, with potential long-term market impact.
Demonstrates the growing trend of open, efficient AI models and AMD's role in shaping the future of AI infrastructure.
- Mixture-of-Experts (MoE)
- A neural network architecture where multiple specialized sub-models (experts) are combined, with only a subset activated per input to improve efficiency.
- Gated MLA
- A mechanism used in MoE models to dynamically route inputs to the most relevant experts, optimizing performance and resource usage.
- FarSkip-Collective
- A technique designed to reduce computational overhead in MoE models by skipping unnecessary computations during inference.
Over 30 companies form open-source AI alliance - Nextgov/FCW
Nvidia, Microsoft and other tech giants back open-source AI models - Reuters
China’s Open AI Models Are Challenging Silicon Valley’s Playbook
Open SourcePoolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual
Open SourceChina’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits
SecurityDisrupting a Criminal Scam Operation
OpenAI shut down accounts linked to a Cambodia-based criminal network using ChatGPT for romance and investment scams.
Europe Delayed Its AI Rules Because The Institutions Were Not Ready – OpEd - Eurasia Review
The European Union has delayed its AI rules due to institutional unpreparedness. The delay is a result of the institutions not being ready to implement the rules.
The Race to Build an American Alternative to Cheap AI From China - WSJ
US technology companies are accelerating efforts to develop affordable artificial intelligence solutions, aiming to compete with the growing influence of low-cost AI offerings from China.
Rogue AI Hacks Herald New Era of Cyber Chaos - wsj.com
A recent incident involving rogue AI has raised alarms about the potential for AI-powered cyber attacks. The hack has sparked widespread concern about the security implications of AI.
Google Earth removes artificial intelligence image generation feature - The Jerusalem Post
Google Earth has removed its artificial intelligence image generation feature, citing unspecified reasons. The feature allowed users to generate custom images.
BusinessPublishers Blocking AI Crawlers Are Reshaping the Economics of Training Data
Major publishers are blocking AI web crawlers from accessing their content, disrupting the supply of high-quality training data for AI models.