Interpreting neural networks for biological sequences by learning stochastic masks

Johannes Linder,Alyssa La Fleur,Zibo Chen,Ajasja Ljubetič,David Baker,Sreeram Kannan,Georg Seelig

NATURE MACHINE INTELLIGENCE（2022）

引用 12|浏览29

暂无评分

摘要

Sequence-based neural networks can learn to make accurate predictions from large biological datasets, but model interpretation remains challenging. Many existing feature attribution methods are optimized for continuous rather than discrete input patterns and assess individual feature importance in isolation, making them ill-suited for interpreting nonlinear interactions in molecular sequences. Here, building on work in computer vision and natural language processing, we developed an approach based on deep learning—scrambler networks—wherein the most important sequence positions are identified with learned input masks. Scramblers learn to predict position-specific scoring matrices where unimportant nucleotides or residues are scrambled by raising their entropy. We apply scramblers to interpret the effects of genetic variants, uncover nonlinear interactions between cis -regulatory elements, explain binding specificity for protein–protein interactions, and identify structural determinants of de novo-designed proteins. We show that scramblers enable efficient attribution across large datasets and result in high-quality explanations, often outperforming state-of-the-art methods.

查看译文

关键词

Genome informatics,Machine learning,Protein function predictions,Protein structure predictions,Engineering,general

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要