About

I am a GSK-funded DPhil student on the Healthcare Data Science CDT at the University of Oxford. I am a member of the Oxford Applied and Theoretical Machine Learning Group (OATML), supervised by Professor Yarin Gal. My previous research sits at the nexus of black-box optimisation and AI-driven scientific discovery: I characterise the hidden topography of optimisation landscapes, and build agentic systems that synthesise scientific knowledge. I believe a deep understanding of the problem itself and existing literature is a fundamental prerequisite for building effective, autonomous scientific discovery systems like self-driving laboratories.

Before Oxford, I completed my MSc in Computer Science at the University of Electronic Science and Technology of China. There I developed GraphFLA, a toolbox for characterising black-box optimisation problems such as those in protein engineering and hyperparameter optimization, and OmniScience, a platform that searches, extracts, verifies, and synthesises knowledge across vast literature corpora. These have led to collaborations with medical experts from GE Healthcare and biologists at the John Innes Centre. I spent the final year of my MSc as a Research Assistant at the University of Exeter under the supervision of Professor Ke Li, and visited the Earlham Institute. Outside research, I enjoy photography and cooking.

News

  • Sep 2026Visiting the Earlham Institute (EI) in Norwich.
  • Jul 2026Presented our work on genomic model interpretability as a spotlight poster at ICML 2026 in Seoul.
  • Feb 2026Awarded a GSK-funded studentship in the EPSRC CDT in Healthcare Data Science. I look forward to joining Oxford this September.
  • Jan 2026FAITH, our work with GE Healthcare on fact-checking LLM-generated medical content, was presented in an oral session at AAAI 2026 in Singapore.
  • Dec 2025Presented our work on augmenting ProteinGym and RNAGym through fitness landscape analysis as a spotlight presentation at NeurIPS 2025 in San Diego.
  • Dec 2025Received a $10,000 research grant from Modal. Thank you for the generous support!
  • Oct 2025Received the National Scholarship of China.
  • Jun 2025Presented our work on analysing software configuration landscapes at ISSTA 2025 in Trondheim, Norway.

Research

Theme 1

Fitness landscape analysis

Deciphering the topography of black-box optimisation landscapes. See more at GraphFLA.

Why landscape ruggedness matters Three paired columns each contain a fitness surface above its sparse-prediction cross-section. Ruggedness increases from left to right. Black illustrative optimization trajectories reach the highest peak on the smooth surface, the secondary peak on the middle surface, and a lower local peak after a winding search on the rugged surface. The trajectories are schematic, not results of a particular solver. Below each surface, its cross-section is compared with a shape-preserving surrogate fitted to the same seven sampling positions. Shaded differences illustrate missed variation between samples. These are illustrative examples of local search and sparse prediction, not a universal performance guarantee. Vertical axes: fitness (objective value). One caption beneath each pair identifies low, moderate, or high ruggedness. One shared legend explains all six plots. Optimization trajectory Sampled data points True fitness Surrogate model fitness (objective value) Low ruggedness fitness (objective value) Moderate ruggedness fitness (objective value) High ruggedness

* equal contribution

Neural Information Processing Systems (NeurIPS)NeurIPS 2025 Spotlight · top 2%

Augmenting Biological Fitness Prediction Benchmarks with Landscape Features from GraphFLA

A Python framework that measures the topography of biological fitness landscapes, turning ProteinGym and RNAGym leaderboard results into explanations of when and why mutation-effect predictors work.

Mingyu Huang, Shasha Zhou, Ke Li · Paper Code

Machine learning models increasingly map biological sequence-fitness landscapes to predict mutational effects. Effective evaluation of these models requires benchmarks curated from empirical data. Despite their impressive scales, existing benchmarks lack topographical information regarding the underlying fitness landscapes, which hampers interpretation and comparison of model performance beyond averaged scores. Here, we introduce GraphFLA, a Python framework that constructs and analyzes fitness landscapes from diverse modalities (DNA, RNA, protein, and beyond.), accommodating datasets up to millions of mutants. GraphFLA calculates 20 biologically relevant features that characterize 4 fundamental aspects of landscape topography. By applying GraphFLA to over 5,300 landscapes from ProteinGym, RNAGym, and CIS-BP, we demonstrate its utility in interpreting and comparing the performance of dozens of fitness prediction models, highlighting factors influencing model accuracy and respective advantages of different models. Additionally, we release 155 combinatorially complete empirical fitness landscapes, encompassing over 2.2 million sequences across various modalities. All the codes and datasets are available at https://github.com/COLA-Laboratory/GraphFLA.

International Symposium on Software Testing and Analysis (ISSTA)ISSTA 2025

Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective

Landscape analysis of 86 million configurations of SQLite, LLVM and Apache yields six findings that together characterise the topography of software configuration landscapes.

Mingyu Huang, Peili Mao, Ke Li · Paper Code

Modern software systems are often highly configurable to tailor varied requirements from diverse stakeholders. Understanding the mapping between configurations and the desired performance attributes plays a fundamental role in advancing the controllability and tuning of the underlying system, yet has long been a dark hole of knowledge due to its black-box nature. While there have been previous efforts in performance analysis for these systems, they analyze the configurations as isolated data points without considering their inherent spatial relationships. This renders them incapable of interrogating many important aspects of the configuration space like local optima. In this work, we advocate a novel perspective to rethink performance analysis—modeling the configuration space as a structured “landscape”. To support this proposition, we utilized GraphFLA, an open-source, graph data mining empowered fitness landscape analysis (FLA) framework. By applying this framework to 86M benchmarked configurations from 32 running workloads of 3 real-world systems, we arrived at 6 main findings, which together constitute a holistic picture of the landscape topography that could have implications on both configuration tuning and performance modeling.

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)KDD 2025

On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory Study

1,500 hyperparameter loss landscapes across 6 models, 63 datasets and 11M+ configurations reveal a shared topography and explain why popular HPO methods succeed.

Mingyu Huang, Ke Li · Paper Code

Previous efforts on hyperparameter optimization (HPO) of machine learning (ML) models predominately focus on algorithmic advances, yet little is known about the topography of the underlying hyperparameter (HP) loss landscape, which plays a fundamental role in governing the search process of HPO. While several works have conducted fitness landscape analysis (FLA) on various ML systems, they are limited to properties of isolated landscape without interrogating the potential structural similarities among landscapes induced on different scenarios. The exploration of such similarities can provide a novel perspective for understanding the mechanism behind modern HPO methods, but has been missing. In this paper, we mapped 1,500 HP loss landscapes of 6 representative ML models on 63 datasets across different fidelity levels, with 11M+ configurations. By conducting exploratory analysis on these landscapes with fine-grained visualizations and dedicated FLA metrics, we observed a similar landscape topography across a wide range of models, datasets, and fidelities, and shed light on the mechanism behind the success of several popular methods in HPO. The artifacts associated with this paper is available at https://github.com/COLA-Laboratory/GraphFLA.

International Joint Conference on Artificial Intelligence (IJCAI)IJCAI 2023

Exploring Structural Similarity in Fitness Landscapes via Graph Data Mining: A Case Study on Number Partitioning Problems

Using local optima networks as a proxy, graph mining shows that instances in neighbouring dimensions share landscape structure and that solvers behave alike on them.

Mingyu Huang, Ke Li · Paper

One of the most common problem-solving heuristics is by analogy. For a given problem, a solver can be viewed as a strategic walk on its fitness landscape. Thus if a solver works for one problem instance, we expect it will also be effective for other instances whose fitness landscapes essentially share structural similarities with each other. However, due to the black-box nature of combinatorial optimization, it is far from trivial to infer such similarity in real-world scenarios. To bridge this gap, by using local optima network as a proxy of fitness landscapes, this paper proposed to leverage graph data mining techniques to conduct qualitative and quantitative analyses to explore the latent topological structural information embedded in those landscapes. In our experiments, we use the number partitioning problem as the case and our empirical results are inspiring to support the overall assumption of the existence of structural similarity between landscapes within neighboring dimensions. Besides, experiments on simulated annealing demonstrate that the performance of a meta-heuristic solver is similar on structurally similar landscapes.

Theme 2

Knowledge discovery & synthesis

Agentic systems for knowledge extraction and reasoning. See a demo for AI for Science research at OmniScience.

From evidence to synthesis The same 188 evidence nodes appear in the same positions in all three panels. The first highlights relevant evidence in green among 600 unrelated grey information points. This depicts retrieval from a wider information pool, without implying exhaustive coverage. The second reveals within-topic relationships and cross-topic links. The third highlights four paths through that network, bringing evidence from different topics to a shared synthesis. Schematic illustration. Finding relevant evidence Structured knowledge Evidence-backed synthesis
International Conference on Machine Learning (ICML)ICML 2026 Spotlight · top 5%

Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods

Systematic literature mining reveals that evaluations of interpretability methods often rely on cherry-picked results, while the methods themselves are often unreliable. We argue for a tiered reporting framework to improve the scientific rigour of explainable AI (xAI).

Shasha Zhou, Mingyu Huang, Ke Li · Paper

Advances in machine learning and computational power have unlocked the predictive potential of the human genome, yet biologists increasingly demand that these models also elucidate the underlying biological mechanisms. While interpretable machine learning (IML) techniques have been increasingly applied to bridge this gap, there has been a pervasive reliance on anecdotal validation: the vast majority of research employs a single IML method and reports only isolated successful instances. Through a benchmarking study on transcription factor binding, we demonstrate the risks of current practices. We show that different IML methods can often (1) yield contradictory explanations for identical predictions, (2) fail to localize known regulatory motifs, and (3) do not faithfully reflect the model’s internal decision process. In light of this, we argue for a validation framework analogous to clinical trials. Just as trials require rigorous design and the reporting of adverse events, genomic interpretability must move beyond cherry-picked plausibility toward systematic assessment of consistency, faithfulness, and biological validity. To facilitate this, we propose a tiered framework to guide the rigorous evaluation and reporting of genomic IML methods.

International Joint Conference on Artificial Intelligence (IJCAI)IJCAI 2025

Conversational Exploration of Literature Landscape with LitChat

A conversational literature agent that retrieves papers, builds and analyses knowledge graphs and topic maps, and turns systematic reviews into conversations grounded in evidence.

Mingyu Huang, Shasha Zhou, Yuxuan Chen, Ke Li · Paper Code

We are living in an era of “big literature”, where the volume of digital scientific publications is growing exponentially. While offering new opportunities, this also poses challenges for understanding literature landscapes, as traditional manual reviewing is no longer feasible. Recent large language models (LLMs) have shown strong capabilities for literature comprehension, yet they are incapable of offering “comprehensive, objective, open and transparent” views desired by systematic reviews due to their limited context windows and trust issues like hallucinations. Here we present LitChat, an end-to-end, interactive and conversational literature agent that augments LLM agents with data-driven discovery tools to facilitate literature exploration. LitChat automatically interprets user queries, retrieves relevant sources, constructs knowledge graphs, and employs diverse data-mining techniques to generate evidence-based insights addressing user needs. We illustrate the effectiveness of LitChat via a case study on AI4Health, highlighting its capacity to quickly navigate the users through large-scale literature landscape with data-based evidence that is otherwise infeasible with traditional means.

Molecular Plant 2026

PlantScience.ai: An LLM-Powered Virtual Scientist for Plant Science

A virtual plant biologist built on an automatically constructed, continuously updated knowledge graph, grounding every answer in traceable primary sources.

Haopeng Yu*, Shasha Zhou*, Mingyu Huang*, Ling Ding, Yuxuan Chen, et al., Yiliang Ding, Ke Li · Paper

The accelerating growth of plant science knowledge presents a major challenge to extracting accurate, up-to-date knowledge from an increasingly fragmented and domain-specific corpus. General-purpose large language models, while powerful, often misinterpret plant science terminology and lack mechanisms for source traceability. We created PlantScience.ai, a virtual plant biology scientist powered by an automated scientific knowledge graph construction pipeline. PlantScience.ai exhibits expert-level reasoning in plant biology and maintains scholarly rigor inciting the literature. Through continuous learning, it integrates the latest research to ensure that its knowledge base remains current and scientifically robust. Apart from providing the answers to scientific questions, PlantScience.ai can interact with human scientists, follow instructions, and retrieve information with citation awareness, grounding each response in primary sources to ensure accuracy and verifiability. PlantScience.ai marks a pivotal advance toward a collaborative scientific paradigm in which virtual and human plant scientists work synergistically to accelerate discovery while preserving the unique value of human insight. PlantScience.ai is available at https://plantscience.ai.

AAAI Conference on Artificial Intelligence (AAAI)AAAI 2026 Oral · top 5%

Assessing Automated Fact-Checking for Medical LLM Responses with Knowledge Graphs

FAITH scores the factuality of medical LLM answers without reference answers by breaking responses into claims and mapping them onto a medical knowledge graph.

Shasha Zhou, Mingyu Huang, Jack Cole, Charles Britton, Ming Yin, Jan Wolber, Ke Li · Paper

The recent proliferation of large language models (LLMs) holds the potential to revolutionize healthcare, with strong capabilities in diverse medical tasks. Yet, deploying LLMs in high-stakes healthcare settings requires rigorous verification and validation to understand any potential harm. This paper investigates the reliability and viability of using medical knowledge graphs (KGs) for the automated factuality evaluation of LLM-generated responses. To ground this investigation, we introduce FAITH, a framework designed to systematically probe the strengths and limitations of this KG-based approach. FAITH operates without reference answers by decomposing responses into atomic claims, linking them to a medical KG, and scoring them based on evidence paths. Experiments on diverse medical tasks with human subjective evaluations demonstrate that KG-grounded evaluation achieves considerably higher correlations with clinician judgments and can effectively distinguish LLMs with varying capabilities. It is also robust to textual variances. The inherent explainability of its scoring can further help users understand and mitigate the limitations of current LLMs. We conclude that while limitations exist, leveraging KGs is a prominent direction for automated factuality assessment in healthcare.

Grants

Research grants

  • 2026EPSRC & GSK CDT Studentship in Healthcare Data Science
  • 2026OpenAI Academic Researcher Program ($12,000)
  • 2026Modal Research Grant ($10,000)
  • 2025Modal Research Grant ($10,000)