Tabular LLMs: An Introduction to the Foundation Models That Predict Your Spreadsheet
Tabular foundation models can predict missing spreadsheet columns zero-shot and have outperformed tuned gradient-boosted trees on the TabArena benchmark.

- Tabular LLMs predict missing data without specific training examples.
- These models have outperformed tuned XGBoost on the TabArena benchmark.
- XGBoost remains competitive in certain specific use cases.
- The shift challenges the dominance of tree-based algorithms in data science.
Tabular foundation models represent a shift in machine learning by applying large language model architectures to structured data. Unlike traditional methods that require task-specific training, these models can predict missing values in spreadsheets zero-shot, similar to how text LLMs complete sentences.
Recent evaluations on the TabArena benchmark indicate that these tabular LLMs have surpassed fully tuned gradient-boosted trees like XGBoost. This performance suggests that pre-trained foundation models may offer superior generalization capabilities for data science tasks compared to the long-standing standard of tree-based algorithms.
The article provides a technical breakdown of how these models function and includes an independent reproduction of the leading open-source options. It also maps out the specific scenarios where XGBoost continues to maintain an advantage over these newer foundation models.
Offers a new approach to feature engineering and prediction tasks without extensive model tuning.
Could improve data accuracy and reduce the time required to prepare predictive models.
Highlights a disruption in the established data science market dominated by gradient boosting.
- Zero-shot
- The ability of a model to perform a task without seeing any specific examples for that task during training.
- Gradient-boosted trees
- A traditional machine learning technique that builds prediction models using an ensemble of decision trees.
New tool identifies the sources of fake video - University of California, Riverside
Weak AI regulations may leave artificial intelligence less safe - Earth.com
Katy ISD restricts the use of AI in elementary school classrooms - Houston Public Media
KV the Apostate: Faith-Based Computing Versus the Unnatural Science - Communications of the ACM
AI ResearchAnthropic releases Opus 5 with ‘close’ to Fable 5’s capabilities
LLMAnthropic's Opus 5 is about token efficiency, not a capability leap
Anthropic's latest model, Opus 5, prioritizes token efficiency to reduce operational costs and improve practical deployment, rather than focusing solely on a significant leap in raw intelligence capabilities. This strategic move addresses the growing demand for more cost-effective large language models.
AI ToolsContext Compression: Making AI Agents Forget Without Losing the Plot
Rijul is developing a micro AI code reviewer called git-lrc, which uses context compression to help AI agents forget unnecessary information.
Katy ISD launches artificial intelligence framework for 2026-27 school year | Katy - Fulshear - Community Impact
Katy ISD has introduced an artificial intelligence framework for the upcoming 2026-27 school year, aiming to enhance student learning experiences.
BusinessRFK Jr.'s hand-picked committee approves manufacture of peptides he uses
A committee led by RFK Jr. has approved the manufacture of peptides he uses, despite a lack of human safety and efficacy data.
LLMAnthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price
Anthropic introduced Claude Opus 5, its latest flagship model, which claims to match Fable 5 performance while costing half as many tokens. The model achieved a 30.2% score on the ARC-AGI-3 benchmark, outpacing GPT‑5.6 Sol.
Nvidia, Microsoft and other tech giants back open-source AI models - Reuters
Nvidia and Microsoft have joined a coalition of major technology companies to support open-source artificial intelligence models. This alliance aims to foster open innovation and safety in AI development.