Weekly BioML Digest [August 24, 2026]
Machine Learning × Computational Biology paper compilation
Hey! It's your weekly digest of machine learning papers in CompBio and Drug Discovery.
Feedback? Email me at biomldigest@gmail.com.
📚 Peer-Reviewed Journals (Top 20)
1120 matched filters -> 20 selected after LLM relevance + novelty ranking.
-
🧫 A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability
Erasmus, M. Frank, Bedinger, Daniel, ..., Ferrara, Fortunato, Bradbury, Andrew R. M. — Nature Biotechnology, 2026-08-19
Prospective, blinded benchmark of AI antibody design shows several methods can produce sub‑100 pM, developable antibodies, but performance does not generalize across tasks, exposing key gaps in affinity prediction and out‑of‑library design. -
🤖 Quantitative and interface-aware prediction of peptide–protein interactions by VITAL
Chen, Wei-Hao, Wang, Qi-Wen, ..., Lin, Chen, Ji, Zhi-Liang — Nature Machine Intelligence, 2026-08-19
Introduces VITAL, a dual‑channel deep model that fuses protein language embeddings with geometry‑aware encoders to predict peptide–protein binding, map interfaces, and estimate affinity, achieving state‑of‑the‑art accuracy with experimental validation. -
📰 Access to all stereoisomers of chiral alcohols with multiple stereocentres enabled by machine learning-empowered protein engineering
Lu, Zhenyu, Zhou, Jiahui, ..., Huang, Meilan, Wu, Qi — Nature Synthesis, 2026-08-21
Machine‑learning‑guided protein engineering evolves an alcohol dehydrogenase into eight stereocomplementary variants, enabling access to all stereoisomers from 2,2‑disubstituted cyclodiketones with up to 99% selectivity. -
📰 GRASSP: RNA Language Model-Enhanced Graph Attention with Adaptive Gating for RNA-Small Molecule Binding Site Prediction.
Thi Lan Nguyen, Nguyen Quoc Khanh Le — Bioinformatics (Oxford, England), 2026-08-22
GRASSP integrates pretrained RNA language model embeddings with adaptive graph attention to predict RNA–small molecule binding sites, outperforming baselines by up to 24% AUC and 45% MCC while reducing structural annotation needs.
Affiliations: School of Electrical Engineering and Computer Science, The University of Queensland; AIBioMed Lab, Taipei Medical University; ... -
📡 Drug screen and machine learning predict neuroprotective agents in a preclinical human model of childhood dementia
Greenberg, Zarina, McDonald, Ella, ..., Smith, Nicholas, Bardy, Cedric — Nature Communications, 2026-08-20
A multimodal platform integrating machine learning with imaging, single‑nuclei transcriptomics, and electrophysiology identifies repurposed neuroprotective compounds in patient‑derived neural models of childhood dementia (MPS IIIA). -
📰 AI-assisted multi-omics and virtual screening identify the IGF2BP2-SGMS2 axis for PDAC precision immunotherapy
Chen, Yutong, Yan, Leye, ..., Zhang, Yu, Zhang, Weiyu — npj Precision Oncology, 2026-08-20
AI‑assisted single‑cell multi‑omics and ML modeling identify the IGF2BP2–SGMS2 axis as a druggable driver of immune evasion in PDAC; virtual screening nominates montelukast to inhibit IGF2BP2 function and synergize with PD‑L1 blockade. -
📰 Data-centric training enables meaningful interaction learning in protein-ligand binding affinity prediction
Silva, Matheus Müller Pereira, Vidal, Lincon Onório, ..., Custódio, Fábio Lima, Dardenne, Laurent Emmanuel — Journal of Cheminformatics, 2026-08-22
Demonstrates data‑centric training (molecular dropout + rotation augmentation) enables a simple CNN to learn true protein–ligand interactions and generalize under stringent Pfam‑CV splits without synthetic data or complex architectures. -
🚀 Large language models enhance annotation of enzymes in metagenomes
Lei Zheng, Bowen Li, Siqi Xu, Junnan Chen, Guanxiang Liang — Science Advances, 2026-08-19
LLM‑driven enzyme annotation and a full metagenomic pipeline improve discovery of enzymatic dark matter with demonstrated clinical‑cohort utility. -
📰 ViralMap: predicting features in viral proteins from primary sequence.
Shrish Dwivedi, Shaunak Kar, Andrew P Horton, Jimmy D Gollihar — Journal of virology, 2026-08-18
ViralMap uses ESM‑2 embeddings to perform multi‑label, residue‑level annotation of viral proteins (10 classes), generalizing to unseen families and aiding antigen engineering for rapid vaccine design.
Affiliations: Systems, Synthetic; Department of Pathology & Genomic Medicine, Antibody Discovery and Accelerated Protein Therapeutics -
📰 QwenCryoMarker: a universal post-processing framework for contamination-aware particle cleaning.
Yunhai Sun, Jiahao Zhao, Nan Xu, Lei Wang, Wei Ding, Ming Li — Acta crystallographica. Section D, Structural biology, 2026-08-20
QwenCryoMarker combines a fine‑tuned vision‑language model with contamination‑aware filtering to clean cryo‑EM particle picks, improving precision/F1 and downstream 3D reconstructions across diverse datasets and pickers. -
📰 Cloud-based ligand-guided virtual screening with deep-learning-enhanced docking identifies a micromolar CXCR4 antagonist
Nalinratana, Nonthaneth, Sangsawat, Monsin, ..., Vajragupta, Opa, Rojsitthisak, Pornchai — Molecular Diversity, 2026-08-20
Cloud‑executed, deep‑learning‑enhanced docking workflow identifies a micromolar CXCR4 antagonist validated by binding, chemotaxis, and patient‑derived organoids, showcasing an accessible early‑stage discovery pipeline. -
📰 To ML-Predict or Not to ML-Predict: The Impact of Machine Learning-Predicted Protein Structures on FEP Accuracy and Data Augmentation
Parker Dryja, Morné Muller, Monique Horn, Ilya A. Balabin, ..., Yuri K. Peterson, F. Joubert, T. Kaiser, P. Burger — Journal of Chemical Information
and Modeling, 2026-08-20
Systematic assessment shows ML‑predicted protein structures can support FEP variably; micro/macro conformational state differences, not structural source, chiefly govern predictive reliability, underscoring the need for careful validation. -
📰 HFGuidedDesign: de novo design of cyclic peptide binders via structure-guided discrete diffusion.
Haomeng Hu, Renjie Zhu, Ning Zhu, Chengyun Zhang, ..., Chongyang Li, Jingjing Guo, Xudong Wang, Hongliang Duan — Chemical science, 2026-08-19
HFGuidedDesign couples discrete diffusion with HighFold structure guidance to design cyclic peptide binders (head‑to‑tail/disulfide) in real time, achieving 67–75% success against two targets.
Affiliations: College of Pharmaceutical Sciences, Zhejiang University of Technology Hangzhou 310014 China xdwang2019@zjut.edu.cn.; Faculty of Applied Sciences, Macao Polytechnic University Macao 999078 China hduan@mpu.edu.mo.; ... -
🏛️ Molecular assembly as a universal biosignature measurable by mass spectrometry.
Lindsay A Rutter, Abhishek Sharma, Ian Seet, David Obeh Alobo, An Goto, Leroy Cronin — Proceedings of the National Academy of Sciences of the United States of America, 2026-08-18
Establishes molecular assembly (MA) as a mass‑spectrometry‑measurable biosignature; an ML model predicts MA from MS1 spectra with 3× lower error than baselines, enabling life detection without structure elucidation.
Affiliations: School of Chemistry, University of Glasgow -
📰 DynMoCo: A novel AI framework to reveal modular substructures of protein from molecular dynamics.
Lingchao Mao, Mingu Kwak, Amir Hossein Kazemipour Ashkezari, Zhenhai Li, ..., Jung Hun Phee, Sooyeon Kang, Jing Li, Cheng Zhu — Biophysical journal, 2026-08-18
DynMoCo is an interpretable graph‑based deep framework that detects dynamic protein communities from MD trajectories, revealing force‑induced substructure reorganization during integrin unbending.
Affiliations: H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology; George W. Woodruff School of Mechanical Engineering, Georgia Institute of Technology; ... -
📰 Arrhythmia risk predictions from molecular simulations of cardiac ion channel-drug interactions.
Kyle C Rouen, Kush Narang, Yanxiao Han, David Wang, ..., Sophia Brunkow, Vladimir Yarov-Yarovoy, Alexander D MacKerell, Igor Vorobyov — Biophysical journal, 2026-08-18
Physics‑based docking (SILCS) with Bayesian ML predicts channel‑drug interactions across hERG/NaV1.5/CaV1.2; potassium channel fragment features (e.g., cationic nitrogen) emerge as key determinants of proarrhythmic risk.
Affiliations: Department of Physiology and Membrane Biology, University of California; Department of Physiology and Membrane Biology, University of California; ... -
📰 An Integrated Consensus Machine Learning and Structure-Based Workflow for the Discovery of Novel Tankyrase 1 Inhibitors
M. Bilotta, Adriana Gargano, R. Rocca, V. Maggisano, S. Bulotta, Stefano Alcaro — Pharmaceuticals, 2026-08-19
Consensus ML screening plus structure‑based refinement and 500‑ns MD identify and validate a potent TNKS1 inhibitor (∼80% inhibition at 0.1 µM), illustrating an integrated AI→SBVS→MD→assay pipeline. -
📰 Algorithmic Determinants of Performance Heterogeneity in Whole-Genome Sequencing-Based Prediction of Drug Resistance in Mycobacterium tuberculosis: A Systematic Review and Meta-Analysis
Baozhen Peng, Yang Zhou, Xiangchen Li, Huihui Liu, Bing Zhao, Ping Hou, X. Ou, Yanlin Zhao — Microorganisms, 2026-08-18
Systematic review/meta‑analysis of WGS‑based TB resistance prediction highlights strong overall specificity and variable sensitivity across tools; algorithmic choices and validation design drive heterogeneity and generalizability. -
📰 Benchmarking the Impact of Data Leakage on the Performance of Knowledge Graph Embedding Models for Biomedical Link Prediction.
Galadriel Bri Ere, Thomas Stosskopf, Benjamin Loire, Anäıs Baudot — Bioinformatics (Oxford, England), 2026-08-20
Benchmarks data leakage in biomedical knowledge graph embeddings, showing train–test redundancies inflate performance and that random/cold‑start splits overestimate real‑world drug repurposing generalization.
Affiliations: Aix Marseille Univ, INSERM; Neurology Therapeutic Area, R&D Servier Paris-Saclay Institut; ... -
📰 Identification of a novel small-molecule modulator targeting SNX10 to inhibit osteoclastic bone resorption.
Yihe Li, Qihang Wu, Shengnan Qin, Ruth Seeber, ..., Scott G Wilson, Alice Vrielink, Haiming Jin, Jiake Xu — Journal of bone and mineral research : the official journal of the American Society for Bone and Mineral Research, 2026-08-23
AI‑enhanced screening discovers an SNX10‑targeting antiresorptive that preserves osteoclastogenesis while blocking resorption, with in vivo efficacy in estrogen‑deficiency bone loss.
Affiliations: School of Biomedical Sciences, The University of Western Australia; School of Molecular Sciences, The University of Western Australia; ...
🧬 Preprints (arXiv + bioRxiv)
51 matched filters -> 20 selected after LLM relevance + novelty ranking.
-
🧬 A Generative Virtual Tissue Model Enables Computational Design of Therapeutic Perturbation Strategies
Lu, Y.; Zhang, W.; Chen, Y.-J.; Yin, J.; ...; Wang, Z. J.; Poon, Y.; You, Y.; Thomson, M. — bioRxiv, 2026-08-20
Introduces a generative geometric GNN (Cell Interaction Foundation Model) trained by masked-transcriptome prediction to simulate tissue transcriptional dynamics from spatial seeds and design single/combinatorial therapeutic perturbations in silico. Validated across gene prediction, disease classification, and recapitulation of perturbation responses, enabling large-scale therapeutic strategy design.
Affiliations: Division of Biology and Biological Engineering, California Institute of Technology -
📄 Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens — arXiv, 2026-08-18
Presents MCTH, an MCTS-based planning layer over frozen folding and inverse-folding models for unified sequence–structure co-design across proteins, RNA/DNA, and ligands. Adaptive search improves design quality under fixed inference budgets and transfers to held-out AlphaFold3/Chai-1 evaluations without fine-tuning. -
📄 FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design
Guofeng Zhang, Rong Han, Xiaoyu Wang, Zhiyun Li, Zongbo Han, Xiaohong Liu, Guangyu Wang — arXiv, 2026-08-20
Proposes FAR-DPO, a feasibility-aware, group-robust preference optimization that steers generative models toward biophysically feasible cyclic peptides. Improves target-wise success and binding scores on CPSea LNR under fixed generation budgets. -
📄 PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening
Jia-Qi Lin, Yinghua Yao, Chang-Dong Wang, Yew-Soon Ong, Yuangang Pan — arXiv, 2026-08-20
Develops PETA, a parameter-efficient test-time adaptation for virtual screening that constructs pocket-specific negatives via molecular diffusion and embedding mixup. By tuning only LayerNorms, it outperforms pretrained and fully retrained baselines on diverse targets. -
📄 ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries
Seungheun Baek, Mogan Gim, Jaewoo Kang — arXiv, 2026-08-21
ReCurveflow learns curved reaction trajectories via flow matching supervised on NEB-derived paths with off-path correction to resist exposure bias. Achieves state-of-the-art transition-state geometry prediction and provides better initializations for NEB optimization. -
🧬 PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring
Panda, P. K. — bioRxiv, 2026-08-20
PandaDock integrates flexible-ligand search with an SE(3)-equivariant GNN affinity scorer trained at scale, plus fast exact grid evaluation. Delivers competitive pose recovery and affinity generalization and is released as a rigorous open-source docking platform.
Affiliations: Stanford University -
📄 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov — arXiv, 2026-08-19
Trains a chemical plausibility-aware LLM (C3LM) on ~45.6M reactions with Top-K prompting and plausibility/novelty rewards for single-step retrosynthesis. Achieves state-of-the-art OOD performance and explores complementary reaction space to conventional models. -
📄 PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon — arXiv, 2026-08-19
Introduces PGFS++, a synthesis-aware RL framework that improves molecular properties while enforcing explicit forward synthesis routes and preserving input-specific diversity. Prevents reward-hacking collapse and yields diverse, synthesizable improvements. -
🧬 CALFP-MHC: Interpretable Pan-Allelic Prediction of Peptide-MHC Binding and Presentation Using Chemically Grounded Fingerprints and Contrastive Learning
Pham, M.-D. N.; Ho, T.-K.-C.; Nguyen, H.-N.; Tran, L.-S.; Phan, M.-D.; Nguyen, V. — bioRxiv, 2026-08-18
CALFP-MHC encodes amino acids with chemically grounded fingerprints and uses supervised contrastive learning with a CNN/transformer backbone to predict pan-allelic peptide–MHC binding. Maintains high AUC under extreme class imbalance with interpretable anchor determinants.
Affiliations: Medical Genetics Institute -
📄 Off-Manifold Collapse in Guided Protein Language Models
Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang — arXiv, 2026-08-19
Identifies off-manifold collapse in guided protein language models and introduces Mahalanobis filtering over activation density to reject degenerate generations. Improves both targeted property scores and structural plausibility without modifying the generator. -
🧬 Structure-enabled enzyme function prediction unveils elusive terpenoid biosynthesis in archaea
Samusevich, R.; Akmese, S. M.; Hebra, T.; Bushuiev, R.; ...; Kampranis, S. C.; Major, D. T.; Sivic, J.; Pluskal, T. — bioRxiv, 2026-08-20
EnzymeExplorer combines alignment-driven structural domain analysis with protein language models for enzyme function prediction. Discovers and validates archaeal terpene synthases, expanding the known distribution of TPS catalysis.
Affiliations: Czech Academy of Sciences, Institute of Organic Chemistry and Biochemistry -
🧬 MemBack: An Equivariant Graph Neural Network for Backmapping Lipid Membranes
Tunc, Y. E.; Böckmann, R. A. — bioRxiv, 2026-08-23
MemBack employs an SE(3)-equivariant GNN to backmap Martini 3 membrane coarse-grained simulations to CHARMM36 heavy-atom structures in a single pass. Preserves bilayer organization and scales to million-atom systems that can be propagated atomistically.
Affiliations: Friedrich-Alexander-Universität Erlangen-Nürnberg -
🧬 Resolution-standardized evaluation of ligand atomic coordinates in crystallographic structures using machine learning
Miyaguchi, I.; Hata, H.; Kuribayashi, T.; Takahashi, S.; ...; Matsumoto, S.; Terayama, K.; Ohta, M.; Ikeguchi, M. — bioRxiv, 2026-08-20
Introduces the atomic Box Correlation Coefficient and trains a 3D‑CNN (QAEmap) to predict atom-wise coordinate–density consistency in a resolution-standardized framework. Enables robust ligand validation up to ~3.5 Å across crystallographic datasets.
Affiliations: Yokohama City University -
🧬 Self-supervised generation of realistic training data enables nanoscale localization in challenging conditions
Goldenberg, O.; Daniel, T.; Xiao, D.; Shalev Ezra, Y.; Alalouf, O.; Shechtman, Y. — bioRxiv, 2026-08-19
A physics-informed generative model with PSF-aware decoding learns directly from experimental microscopy to synthesize fully labeled, realistic training data. Significantly boosts downstream nanoscale emitter localization under complex backgrounds and low SNR.
Affiliations: Technion Israel Institute of Technology -
📄 Conformal Prediction for Molecular Properties under Label Shift
Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin — arXiv, 2026-08-18
Proposes conformal prediction tailored to label shift by weighting scores with marginal label ratios to produce valid predictive intervals without retraining. Enhances reliability and regulatory-aligned uncertainty quantification for molecular property prediction. -
🧬 An AI System for Autonomous Algorithm Evolution in Drug Development
Zhou, Z.; Nan, Y.; Mou, M.; Qian, Y.; ...; Mi, T.; Sun, H.; Liu, P.; Zhu, F. — bioRxiv, 2026-08-20
DrugEvolve is a multi-role LLM system that autonomously designs, implements, evaluates, and refines algorithms across 11 drug-development tasks. Improves performance on 120 benchmarks and generalizes across sequences, graphs, topology, and text.
Affiliations: Zhejiang University -
🧬 Integration of proteomic data from cell lines and tumors
Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmüller, U.; Raue, A. — bioRxiv, 2026-08-19
ProtInt leverages deep learning with integrated imputation to align proteomes from cell lines and treatment-naive tumors. Improves cross-domain comparability and highlights adapted pathways, aiding selection of clinically relevant cancer models.
Affiliations: University of Augsburg -
📄 Multitask Bayesian Neural Networks for Multiparameter Protein Engineering
Fabio Herrera-Rocha, David Medina-Ortiz, Desiree Wyrzykala, Tharun Srinivasan Sudha, Mehdi D. Davari — arXiv, 2026-08-19
Benchmarks multitask Bayesian neural networks for simultaneous protein property prediction across 27 datasets, identifying Bayesian last-layer models with dimensionality reduction as robust and well-calibrated. Provides practical guidance for data-scarce, multiparameter protein engineering. -
🧬 LAMBDA: A Prophage Detection Benchmark for Genomic Language Models
Lindsey, L. M.; Pershing, N. L.; Dufault-Thompson, K.; Gwak, H.-j.; ...; Stephens, W. Z.; Blaschke, A. J.; Sundar, H.; Jiang, X. — bioRxiv, 2026-08-17
LAMBDA benchmarks genomic language model embeddings for prophage detection via probing, fine-tuning, diagnostics, and genome-wide tasks. Reveals the importance of domain-specific training and current limits of DNA LMs for sequence-level genome annotation.
Affiliations: National Institutes of Health -
📄 Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff — arXiv, 2026-08-18
Shows that domain adaptation of molecular language model encoders to target make-on-demand libraries consistently improves discovery and sample efficiency across drug, materials, and catalysis screens. Adapted encoders often surpass native embeddings and rival strong fingerprint baselines.