Weekly BioML Digest [August 24, 2026]

Share
Weekly BioML Digest [August 24, 2026]

Machine Learning × Computational Biology paper compilation

Hey! It's your weekly digest of machine learning papers in CompBio and Drug Discovery.

Feedback? Email me at biomldigest@gmail.com.

📚 Peer-Reviewed Journals (Top 20)

1120 matched filters -> 20 selected after LLM relevance + novelty ranking.

🧬 Preprints (arXiv + bioRxiv)

51 matched filters -> 20 selected after LLM relevance + novelty ranking.

  • 🧬 A Generative Virtual Tissue Model Enables Computational Design of Therapeutic Perturbation Strategies
    Lu, Y.; Zhang, W.; Chen, Y.-J.; Yin, J.; ...; Wang, Z. J.; Poon, Y.; You, Y.; Thomson, M. — bioRxiv, 2026-08-20
    Introduces a generative geometric GNN (Cell Interaction Foundation Model) trained by masked-transcriptome prediction to simulate tissue transcriptional dynamics from spatial seeds and design single/combinatorial therapeutic perturbations in silico. Validated across gene prediction, disease classification, and recapitulation of perturbation responses, enabling large-scale therapeutic strategy design.
    Affiliations: Division of Biology and Biological Engineering, California Institute of Technology

  • 📄 Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
    Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens — arXiv, 2026-08-18
    Presents MCTH, an MCTS-based planning layer over frozen folding and inverse-folding models for unified sequence–structure co-design across proteins, RNA/DNA, and ligands. Adaptive search improves design quality under fixed inference budgets and transfers to held-out AlphaFold3/Chai-1 evaluations without fine-tuning.

  • 📄 FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design
    Guofeng Zhang, Rong Han, Xiaoyu Wang, Zhiyun Li, Zongbo Han, Xiaohong Liu, Guangyu Wang — arXiv, 2026-08-20
    Proposes FAR-DPO, a feasibility-aware, group-robust preference optimization that steers generative models toward biophysically feasible cyclic peptides. Improves target-wise success and binding scores on CPSea LNR under fixed generation budgets.

  • 📄 PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening
    Jia-Qi Lin, Yinghua Yao, Chang-Dong Wang, Yew-Soon Ong, Yuangang Pan — arXiv, 2026-08-20
    Develops PETA, a parameter-efficient test-time adaptation for virtual screening that constructs pocket-specific negatives via molecular diffusion and embedding mixup. By tuning only LayerNorms, it outperforms pretrained and fully retrained baselines on diverse targets.

  • 📄 ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries
    Seungheun Baek, Mogan Gim, Jaewoo Kang — arXiv, 2026-08-21
    ReCurveflow learns curved reaction trajectories via flow matching supervised on NEB-derived paths with off-path correction to resist exposure bias. Achieves state-of-the-art transition-state geometry prediction and provides better initializations for NEB optimization.

  • 🧬 PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring
    Panda, P. K. — bioRxiv, 2026-08-20
    PandaDock integrates flexible-ligand search with an SE(3)-equivariant GNN affinity scorer trained at scale, plus fast exact grid evaluation. Delivers competitive pose recovery and affinity generalization and is released as a rigorous open-source docking platform.
    Affiliations: Stanford University

  • 📄 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
    Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov — arXiv, 2026-08-19
    Trains a chemical plausibility-aware LLM (C3LM) on ~45.6M reactions with Top-K prompting and plausibility/novelty rewards for single-step retrosynthesis. Achieves state-of-the-art OOD performance and explores complementary reaction space to conventional models.

  • 📄 PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
    Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon — arXiv, 2026-08-19
    Introduces PGFS++, a synthesis-aware RL framework that improves molecular properties while enforcing explicit forward synthesis routes and preserving input-specific diversity. Prevents reward-hacking collapse and yields diverse, synthesizable improvements.

  • 🧬 CALFP-MHC: Interpretable Pan-Allelic Prediction of Peptide-MHC Binding and Presentation Using Chemically Grounded Fingerprints and Contrastive Learning
    Pham, M.-D. N.; Ho, T.-K.-C.; Nguyen, H.-N.; Tran, L.-S.; Phan, M.-D.; Nguyen, V. — bioRxiv, 2026-08-18
    CALFP-MHC encodes amino acids with chemically grounded fingerprints and uses supervised contrastive learning with a CNN/transformer backbone to predict pan-allelic peptide–MHC binding. Maintains high AUC under extreme class imbalance with interpretable anchor determinants.
    Affiliations: Medical Genetics Institute

  • 📄 Off-Manifold Collapse in Guided Protein Language Models
    Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang — arXiv, 2026-08-19
    Identifies off-manifold collapse in guided protein language models and introduces Mahalanobis filtering over activation density to reject degenerate generations. Improves both targeted property scores and structural plausibility without modifying the generator.

  • 🧬 Structure-enabled enzyme function prediction unveils elusive terpenoid biosynthesis in archaea
    Samusevich, R.; Akmese, S. M.; Hebra, T.; Bushuiev, R.; ...; Kampranis, S. C.; Major, D. T.; Sivic, J.; Pluskal, T. — bioRxiv, 2026-08-20
    EnzymeExplorer combines alignment-driven structural domain analysis with protein language models for enzyme function prediction. Discovers and validates archaeal terpene synthases, expanding the known distribution of TPS catalysis.
    Affiliations: Czech Academy of Sciences, Institute of Organic Chemistry and Biochemistry

  • 🧬 MemBack: An Equivariant Graph Neural Network for Backmapping Lipid Membranes
    Tunc, Y. E.; Böckmann, R. A. — bioRxiv, 2026-08-23
    MemBack employs an SE(3)-equivariant GNN to backmap Martini 3 membrane coarse-grained simulations to CHARMM36 heavy-atom structures in a single pass. Preserves bilayer organization and scales to million-atom systems that can be propagated atomistically.
    Affiliations: Friedrich-Alexander-Universität Erlangen-Nürnberg

  • 🧬 Resolution-standardized evaluation of ligand atomic coordinates in crystallographic structures using machine learning
    Miyaguchi, I.; Hata, H.; Kuribayashi, T.; Takahashi, S.; ...; Matsumoto, S.; Terayama, K.; Ohta, M.; Ikeguchi, M. — bioRxiv, 2026-08-20
    Introduces the atomic Box Correlation Coefficient and trains a 3D‑CNN (QAEmap) to predict atom-wise coordinate–density consistency in a resolution-standardized framework. Enables robust ligand validation up to ~3.5 Å across crystallographic datasets.
    Affiliations: Yokohama City University

  • 🧬 Self-supervised generation of realistic training data enables nanoscale localization in challenging conditions
    Goldenberg, O.; Daniel, T.; Xiao, D.; Shalev Ezra, Y.; Alalouf, O.; Shechtman, Y. — bioRxiv, 2026-08-19
    A physics-informed generative model with PSF-aware decoding learns directly from experimental microscopy to synthesize fully labeled, realistic training data. Significantly boosts downstream nanoscale emitter localization under complex backgrounds and low SNR.
    Affiliations: Technion Israel Institute of Technology

  • 📄 Conformal Prediction for Molecular Properties under Label Shift
    Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin — arXiv, 2026-08-18
    Proposes conformal prediction tailored to label shift by weighting scores with marginal label ratios to produce valid predictive intervals without retraining. Enhances reliability and regulatory-aligned uncertainty quantification for molecular property prediction.

  • 🧬 An AI System for Autonomous Algorithm Evolution in Drug Development
    Zhou, Z.; Nan, Y.; Mou, M.; Qian, Y.; ...; Mi, T.; Sun, H.; Liu, P.; Zhu, F. — bioRxiv, 2026-08-20
    DrugEvolve is a multi-role LLM system that autonomously designs, implements, evaluates, and refines algorithms across 11 drug-development tasks. Improves performance on 120 benchmarks and generalizes across sequences, graphs, topology, and text.
    Affiliations: Zhejiang University

  • 🧬 Integration of proteomic data from cell lines and tumors
    Ta, C. Q.; Auth, J. M.; Schilling, M.; Klingmüller, U.; Raue, A. — bioRxiv, 2026-08-19
    ProtInt leverages deep learning with integrated imputation to align proteomes from cell lines and treatment-naive tumors. Improves cross-domain comparability and highlights adapted pathways, aiding selection of clinically relevant cancer models.
    Affiliations: University of Augsburg

  • 📄 Multitask Bayesian Neural Networks for Multiparameter Protein Engineering
    Fabio Herrera-Rocha, David Medina-Ortiz, Desiree Wyrzykala, Tharun Srinivasan Sudha, Mehdi D. Davari — arXiv, 2026-08-19
    Benchmarks multitask Bayesian neural networks for simultaneous protein property prediction across 27 datasets, identifying Bayesian last-layer models with dimensionality reduction as robust and well-calibrated. Provides practical guidance for data-scarce, multiparameter protein engineering.

  • 🧬 LAMBDA: A Prophage Detection Benchmark for Genomic Language Models
    Lindsey, L. M.; Pershing, N. L.; Dufault-Thompson, K.; Gwak, H.-j.; ...; Stephens, W. Z.; Blaschke, A. J.; Sundar, H.; Jiang, X. — bioRxiv, 2026-08-17
    LAMBDA benchmarks genomic language model embeddings for prophage detection via probing, fine-tuning, diagnostics, and genome-wide tasks. Reveals the importance of domain-specific training and current limits of DNA LMs for sequence-level genome annotation.
    Affiliations: National Institutes of Health

  • 📄 Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
    Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff — arXiv, 2026-08-18
    Shows that domain adaptation of molecular language model encoders to target make-on-demand libraries consistently improves discovery and sample efficiency across drug, materials, and catalysis screens. Adapted encoders often surpass native embeddings and rival strong fingerprint baselines.

Read more