PISA maps DNA model insights, removes bias
Reported by Unite.AI: AI Method Reveals What Genomic Models Learn From DNA and Exposes Hidden Experimental Bias. Analysis and context written by TickrWire.
Researchers at the Stowers Institute introduced PISA, a pairwise influence by sequence attribution method that visualizes, at single‑base resolution, what deep‑learning models learn from DNA and can strip experimental bias from MNase‑seq data.

- PISA creates single‑base resolution influence maps that expose both biological signals and experimental bias in genomic AI models.
- By mathematically extracting the MNase‑seq enzyme bias, the authors produced a purified model that learns only true nucleosome biology.
- The bias‑free model identified asymmetric DNA motifs and thousands of chromatin domain boundaries, surpassing some 3D mapping methods.
- Synthetic DNA sequences designed using the model’s rules successfully directed nucleosome placement in wet‑lab tests.
- PISA’s framework is already being re‑implemented by external groups, suggesting rapid community uptake.
On August 25, 2026, the Stowers Institute for Medical Research announced a novel interpretation framework called PISA (pairwise influence by sequence attribution). Led by Julia Zeitlinger and developed with Anshul Kundaje’s group at Stanford, the method allows researchers to trace the contribution of every nucleotide in a DNA sequence to a deep‑learning model’s prediction at a specific genomic location. By generating a two‑dimensional influence map at single‑base resolution, PISA reveals exactly which parts of the input the model relied on, rather than merely providing a single aggregate importance score.
PISA operates inside BPReveal, an extension of the BPNet deep‑learning architecture first released in 2021. The core idea is to compute pairwise influence scores between a target base and all other bases in the surrounding region, producing a matrix that visualizes how each nucleotide affects the model’s output. When applied to MNase‑seq data, a common assay that maps nucleosome positions by cutting exposed DNA while protecting nucleosome‑bound DNA, the method uncovered a distinct fingerprint corresponding to the enzyme’s sequence preference. By mathematically isolating this fingerprint and training a separate bias model, the researchers could subtract the experimental artifact, leaving a purified model that captures only the underlying biological signal.
Sequence‑to‑function neural networks have become a staple in genomics, taking raw DNA as input and predicting outcomes such as transcription‑factor binding or nucleosome organization. However, these models traditionally act as black boxes, offering little insight into why a particular prediction was made. Existing attribution tools collapse each base’s influence into a single number, which can cause positive and negative effects to cancel out and disappear from the analysis. PISA’s high‑resolution approach preserves the full spectrum of influences, allowing the enzyme‑induced bias to emerge clearly as a pattern across the map.
The ability to separate bias from biology mirrors the leap from conventional microscopy to super‑resolution techniques, a comparison made by Zeitlinger in the institute’s press release. Earlier interpretation methods opened the black box; PISA adds enough “pixels” to reveal finer details. In the broader landscape, Google DeepMind’s AlphaGenome, also published in 2026, focuses on predicting regulatory variant effects across the genome. PISA tackles the complementary challenge of interpreting those predictions, showing where the model’s decisions originate. The method has already been re‑implemented by external collaborators and adopted for unrelated biological questions, indicating rapid community uptake.
Beyond bias correction, the bias‑free model uncovered DNA motifs that influence nucleosome positioning over hundreds of base pairs, many of which displayed asymmetry, different effects on the left versus right side of a nucleosome. This asymmetry guided the team to identify chromatin domain boundaries, regions that separate regulatory neighborhoods. Remarkably, the model inferred thousands of such boundaries from MNase‑seq data alone, often with higher precision than traditional three‑dimensional chromatin mapping techniques that require deep sequencing. To test the practical utility of the learned rules, the researchers designed synthetic DNA sequences predicted to arrange nucleosomes in predefined patterns. Experimental validation of a subset of these designs confirmed the model’s predictions, demonstrating that the extracted sequence rules can generate testable hypotheses.
The study also highlights several constraints. The bias‑removal demonstration is specific to MNase‑seq, and while the authors show PISA applied to other assay types, the effectiveness of bias correction may vary. Only a limited number of synthetic sequences were experimentally tested, leaving open questions about scalability. Moreover, the work underscores a persistent gap in the field: the need for teams that combine deep‑learning expertise with hands‑on experimental biology, a combination the authors note as essential for translating computational insights into wet‑lab breakthroughs. Finally, while the method clarifies regulatory DNA mechanisms, its immediate disease relevance remains speculative, as most disease‑associated variants lie in regulatory regions but require additional functional validation.
Looking ahead, PISA provides a template for auditing and refining genomic AI models across diverse datasets. Its ability to isolate experimental artifacts could improve the reliability of large‑scale predictive models, such as those used for variant effect prediction in personalized medicine. The authors anticipate broader adoption of the technique, especially as more labs integrate deep‑learning pipelines into their workflows. Future research may extend PISA to other high‑throughput assays, explore automated bias‑correction pipelines, and combine the approach with emerging large‑scale language models for DNA to deepen our understanding of genome regulation.
Provides a high‑resolution tool for interpreting deep‑learning models on genomic data, enabling more reliable model debugging.
Improves the trustworthiness of AI‑driven genomics pipelines, potentially accelerating drug target discovery and synthetic biology applications.
Demonstrates a tangible advance in AI‑enabled biotech, indicating growing value in companies that combine deep learning with experimental validation.
Offers a concrete example of how model interpretability can be applied to real biological problems, useful for interdisciplinary training.
Shows how AI can uncover hidden biases in scientific data and help generate testable biological hypotheses.
- MNase-seq
- A sequencing assay that maps nucleosome positions by cutting exposed DNA while protecting DNA wrapped around histones.
- Nucleosome
- The basic unit of chromatin, consisting of DNA wrapped around a histone protein core.
- Chromatin domain boundary
- A genomic region that separates distinct regulatory neighborhoods, influencing which genes can interact with which enhancers.
AI bias estimate: The source emphasizes the method’s promise but does not provide extensive validation across multiple assay types, which may overstate its generality. (Automated estimate, not a definitive judgement.)
AI ResearchPew study confirms sharp rise of AI-written text on the web since ChatGPT's launch
AI ResearchKids outlearn AI—and we still don’t know why
AI ResearchWho’s behind the new ‘stealth model’ Ox Alpha?
AI ResearchHarvey Introduces Harvey Tenet: A Kimi K3 Base Post-Trained with Fireworks for Long-Horizon Legal Agent Work
AI ResearchAI could make scientists do more work less well, not less work better, study argues
FundingIndia’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call
Indian voice AI startup Ringg has raised $10 million in a Series A extension led by Peak XV Partners, bringing its total funding in the round to $15.5 million.
RoboticsRobotics startup Generalist reaches $3B valuation, sources say
Robotics startup Generalist secured a nearly $200 million funding extension led by 8VC, lifting its valuation to $3 billion just months after a major Series B round.
BusinessOpenAI loses a top data center exec as stream of high-profile departures continues
OpenAI’s head of data centers, Chris Malone, has left the company as part of a broader executive exodus, raising questions about leadership stability ahead of a planned IPO.
AI ToolsPerplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps
Perplexity introduced Portable Computer, a bundled local‑first AI agent system that runs on NVIDIA DGX Spark and eliminates per‑token fees for on‑device processing.
FundingStability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding
Stability AI announced a $76 million Series B round, bringing its total funding to $232 million. Investors include Universal Music Group, Sony Music, Warner Music, Electronic Arts, AMD Ventures and Pacific Alliance Ventures.
SecurityRussia used ChatGPT to run a covert influence campaign pushing pro-Kremlin narratives across the West
OpenAI banned 36 ChatGPT accounts linked to a Russian influence campaign that used AI to generate pro-Kremlin content across Western platforms.