LFM2.5 230M Runs in-Browser at 1,400 Tokens/sec via WebGPU
Reported by the original publisher: LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels. Analysis and context written by TickrWire.
A 230M-parameter LFM2.5 model runs locally in-browser at 1,400 tokens/sec using custom WebGPU kernels, leveraging prior work from Fable 5 and Opus 4.8.

- LiquidAI/LFM2.5-230M runs locally in-browser at 1,400 tokens/sec using custom WebGPU kernels.
- Kernels were adapted from Fable 5 (shut down) and Opus 4.8, enabling efficient on-device inference.
- Demo available on Hugging Face Spaces for public testing.
- Performance achieved on an M4 Max Mac, demonstrating feasibility on consumer hardware.
- Showcases WebGPU as a viable path for high-performance, client-side AI without dedicated GPUs.
A developer demonstrated the LiquidAI/LFM2.5-230M model running entirely in a web browser via custom WebGPU kernels, achieving 1,400 tokens per second on an M4 Max Mac. The implementation builds on kernels originally developed for Fable 5 (before its shutdown) and Opus 4.8, showcasing efficient on-device inference. A Hugging Face Space provides a live demo for testing. The breakthrough highlights the potential of WebGPU for high-performance, client-side AI workloads without requiring dedicated hardware.
Demonstrates practical WebGPU-based inference for LLMs, reducing dependency on server-side hardware and enabling edge AI applications.
Opens opportunities for privacy-focused, low-latency AI products that run entirely in-browser, reducing cloud costs.
Highlights advancements in on-device AI, potentially disrupting cloud-based inference markets with more efficient alternatives.
Provides a tangible example of WebGPU's capabilities for AI workloads, useful for learning and experimentation.
Shows that advanced AI can run locally on consumer devices, enhancing privacy and accessibility.
- WebGPU
- A modern graphics and compute API for web browsers, enabling GPU acceleration for JavaScript applications.
- Tokens/sec
- A metric measuring the speed of a language model's inference, indicating how many tokens it can process per second.
- GGUF
- A file format for quantized large language models, optimized for efficient inference on consumer hardware.
- Kernel
- A low-level function that performs a specific computation, often optimized for hardware acceleration.
AI bias estimate: Neutral technical demonstration; no overt bias detected. (Automated estimate, not a definitive judgement.)
AI ToolsMeta AI’s new Mac app wants you to talk to your apps
How one Philadelphia school is using AI to strengthen student learning, not replace teachers - CBS News
Domain and publish date filters for Web Search on AgentCore - Amazon Web Services (AWS)
KnowledgeForge: mining gold from the ITSM ticket graveyard - Amazon Web Services (AWS)
Google launches new study tools for Students across Search and Gemini
New White House strategy clarifies military tech priorities: undersea, outer space and AI - Breaking Defense
The White House released a new strategy prioritizing military investments in artificial intelligence, space systems and undersea technologies to counter emerging threats.
AI in an iron grip: How dictatorships use artificial intelligence to strengthen their rule - theins.press
A new report examines how authoritarian governments deploy AI for surveillance, censorship, and propaganda to reinforce their power.
Stripe, OpenRouter finally strike a deal - Banking Dive
Stripe and OpenRouter have partnered to integrate Stripe's payment processing with OpenRouter's AI model aggregation platform.
Exclusive-How a Texas student blew the whistle on a rogue AI hacking attempt - The Mighty 790 KFGO
A Texas student uncovered an AI-powered hacking attempt targeting local systems, prompting a swift law enforcement response.
Student Journalists: AI Is Changing Our Work — And Not For the Better - The 74
A student journalism outlet argues that AI tools are degrading the quality and authenticity of their reporting.
Don’t mistake chatbot intelligence for consciousness - The Economist
The Economist argues that advanced chatbots lack true consciousness despite their impressive intelligence, urging caution against anthropomorphizing AI.