Odysseus Kaloidas

Research vision

AI Agents for Scientific Discovery

I believe AI agents will radically transform how research is conducted over the next decade: enabling investigations across disciplines, accelerating cycles of hypothesis and experiment, and expanding the scale of scientific work. The opportunity is not only to automate existing tasks, but to make new kinds of research possible.

My research focuses on the reasoning, decision-making and learning capabilities that agents need to conduct sustained scientific investigations: generating hypotheses, designing and interpreting experiments, using tools and models of the world, learning from evidence, and deciding what to do next.

My long-term ambition extends from AI co-scientists to autonomous research teams and research factories: systems that sustain investigations over long horizons, learn from success and failure, and produce rigorously validated discoveries. I am particularly interested in applying these ideas to biomedical and therapeutic discovery.

Research interests

Long-horizon scientific agents

How can agents conduct coherent scientific investigations over extended periods of time? I study planning, reasoning, memory, tool use and multi-agent coordination, as well as how agents allocate computation and decide when to reason, retrieve information, use a tool, delegate or stop.

Hypothesis generation, experimentation and verification

How should agents generate hypotheses and determine what evidence is needed to distinguish between them? I study hypothesis search, experiment selection and scientific verification, including the use of predictive and world models to simulate interventions and guide decisions before expensive real-world experiments.

Closed-loop learning and self-improving agents

How can scientific agents improve through experience? I study how experimental outcomes—including failures—should update hypotheses, memory, reasoning strategies and future actions, and how agents can improve their own research processes without reinforcing errors or drifting away from evidence.

Computational foundations and evaluation of scientific discovery

What makes scientific discovery computationally possible, and how should progress be measured? I study the structure and search of hypothesis spaces, the relationship between evidence and scientific value, and the computational, data and experimental resources required for discovery. I am also interested in evaluating whether agentic systems produce genuinely novel, well-supported and useful scientific results.

Example research questions;

  • How can scientific agents reason, remember and act coherently over long research horizons?
  • How should an agent decide whether to reason further, retrieve information, use a tool, consult a world model, run an experiment, collaborate with another agent or stop?
  • How should agents generate and maintain competing hypotheses, and what evidence is required to distinguish between them?
  • How should successful and failed experiments update an agent's hypotheses, memory, reasoning strategies and future research decisions?
  • Can scientific agents improve their own research processes over time, and how can we distinguish genuine self-improvement from accumulated error?
  • How should we evaluate the novelty, evidential support and scientific value of discoveries produced by AI agents?
  • What formal theories and mathematical frameworks can provide a computational foundation for scientific discovery?

Research papers

My research papers examine what scientific agents contribute, how their ideas should be evaluated and what governs the complexity of discovery, through controlled comparisons, model diagnostics and computational analysis. Working with limited compute made careful scoping central to these papers: identifying novel, promising ideas and designing informative experiments that were feasible within the available budget.

The current manuscripts may include revisions made after acceptance to improve readability and clarity.

Do Repeated LLM Decisions Improve CRISPR Hit Discovery?

Odysseus Kaloidas

NeurIPS Workshop · AgenticLS (Spotlight)

When does an AI agent add value to experimental decision-making? Across retrospective CRISPR screens, experimental feedback improves LLM-guided hit discovery, but a simple numerical policy using the same feedback finds more hits on average. Results vary by screen and initialization. The study separates the value of feedback from the value of repeated LLM decisions, providing a controlled way to test where scientific agents genuinely help.

Discarding Answer Ranks Changes the Case for Mixing Biomedical LLMs

Odysseus Kaloidas

NeurIPS Workshop · AgenticLS (Poster)

Combining biomedical LLMs can appear beneficial because of how their answers are counted, not simply because the models are diverse. Holding responses fixed reveals that discarding answer ranks can weaken the single-model comparison and change the apparent benefit of mixing. Under equal call budgets, a selected single model outperforms a mixed-model policy on MedXpertQA. The findings show why model allocation and answer aggregation must be evaluated together.

Position: Toward a Computational Theory of Scientific Discovery

Odysseus Kaloidas

NeurIPS Workshop · AgenticLS (Poster)

Which scientific problems become achievable as AI improves, and when do new capabilities lead to validated discoveries? This position paper proposes a computational framework linking a discovery's resource and information requirements to available methods and the time needed for validation. Mathematical examples and retrospective tests expose why more compute or better benchmark performance need not translate directly into scientific progress, outlining a research program rather than a validated forecast of breakthroughs.

What Does a Protein Surface Add? Ablations Cannot Tell You

Odysseus Kaloidas

NeurIPS Workshop · ICBINB-BIO (Poster)

Is a protein-design model learning from surface geometry, or from information about the sequence it is meant to predict? Audits of two released pipelines identify surface features derived from the target sequence. In a released SurfPro model, changing those features redirects predictions toward a specified alternative sequence. The paper develops tests that distinguish useful representations from target-information leakage, while leaving open how much genuinely backbone-derived surfaces help protein design.

Same Algorithms, Different Scientist: Design Space Bias in AI Discovery

Odysseus Kaloidas

NeurIPS Workshop · AI4MetaScience (Poster)

An AI scientist's apparent contribution can change when its menu of methods is rewritten, even without adding a new executable strategy. Experiments in protein landscapes and knapsack optimization show that ordering and duplicating options alter choices and can reverse conclusions about the agent's added value. Grouping methods by their behavior on fixed tests makes evaluation more stable, without establishing a general performance gain: a step toward measuring scientific judgment rather than menu design.

The Selection-Target Problem in AI-Assisted Science: Reviewer Novelty and Later Scientific Uptake

Odysseus Kaloidas

NeurIPS Workshop · AI4MetaScience (Poster)

What should an AI system learn to value when selecting research ideas? An audit of 3,443 ICLR submissions finds that technical and empirical novelty/significance scores have different relationships with later citation uptake. At a fixed reading budget, the empirical score selects more high-uptake papers in the studied cohort; a separate analysis bounds the effect of missing outcomes. The results motivate testing selection criteria against their intended consequences, without equating citations with scientific value.

Is the Endpoint Well-Posed? Measurement Contracts for AI-Assisted Scientific Discovery

Odysseus Kaloidas

NeurIPS Workshop · AI4MetaScience (Poster)

Does a high discovery-benchmark score reflect scientific reasoning, or recovery of a published conclusion already present in the input? Audits show how revealing or swapping source material changes the recoverability of reference conclusions, and why matching those conclusions does not establish scientific validity. The paper proposes measurement contracts that specify the available evidence, target answer, scoring rule and intended scientific claim, making clearer what a benchmark can actually demonstrate.

When Are Recursive Language Models Useful? A Cost-Aware, Task-Conditional Evaluation

Odysseus Kaloidas

NeurIPS Workshop · LCFM

When does letting a language model write programs and recursively call other models justify the extra cost? Where the required operation is known, task-matched alternatives generally match or outperform recursive execution more cheaply. In open search, advantages over stronger alternatives do not hold consistently across budget-matched repeats. Counting every attempted query, including failures, can reverse apparent winners, offering a practical framework for judging when flexible agent architectures are worth their cost.

Essays

I wrote these essays after noticing that discussions of AI, including talks and writing by leading AI researchers, sometimes blur important conceptual distinctions. The underlying ideas may be understood, yet the way they are communicated can obscure what follows from them. These essays make those distinctions explicit and examine their consequences for reasoning, learning, evaluation and scientific progress.

Lectures

I created these lecture series because I have not found any public ones dedicated to formalizing scientific discovery and connecting it with the latest methods in advanced reasoning and AI agents in the depth and combination I wanted to study. They bring together two complementary questions: how scientific agents should reason, and what a computational theory of discovery should explain.

In addition, even excellent courses I encountered often explained ideas without showing how they arose. These lectures aim to reconstruct that creative process, from the initial question to the design choices and evidence, through three questions:

  1. Why think in this direction?

    What observation, opportunity or dissatisfaction starts the inquiry? Which assumption becomes questionable, and what analogy or reframing makes a new direction conceivable?

  2. How arrive at this particular idea?

    How are candidate solutions constructed? Why pursue one design rather than credible alternatives, and what information could change that choice?

  3. How does it work, and why does it help?

    What mechanism or derivation explains the method? What evidence supports it, which assumptions does it depend on, and where does it fail?

Even strong lectures often focus only on the third question. My aim is to emphasize the first two as well. Explaining an idea is not the same as explaining the creative process behind it. Funnily enough, AI models will probably learn to innovate more effectively this way too.

Opening slide from Reasoning & AI Agents for Scientific DiscoveryReasoning & AI Agents for Scientific Discovery18 lectures 2026An 18-lecture synthesis of advanced reasoning and AI agent methods for scientific discovery: search and test-time computation, latent and diffusion reasoning, hierarchy, reinforcement learning, causal reasoning, memory and multi-agent systems. Worked examples and failure analyses connect these methods to tool use and autonomous experimentation.
Reasoning & AI Agents for Scientific Discovery: lecture topics and slides
No.LectureFocusSlides
01Scientific discovery as a computational processDiscovery as a computational process, with an explicit account of the evidence needed to support a claim.Discovery as a computational process, with an explicit account of the evidence needed to support a claim.PDF
02Patterns of discovery: analogy, recombination, anomaly, mechanismAnalogy, recombination, anomalies and mechanisms as routes to discovery and reasons to change a representation.Analogy, recombination, anomalies and mechanisms as routes to discovery and reasons to change a representation.PDF
03Hypothesis generation and conceptual noveltyConstructing competing causal hypotheses, distinguishing verbal from predictive diversity and choosing interventions that separate them.Constructing competing causal hypotheses, distinguishing verbal from predictive diversity and choosing interventions that separate them.PDF
04Searching hypothesis spacesSearching candidate hypotheses with external verification and feedback, including the example of AlphaEvolve.Searching candidate hypotheses with external verification and feedback, including the example of AlphaEvolve.PDF
05Experiment selection and information gainChoosing experiments by information gain and decision value, including when biased simulations are worth less than direct measurements.Choosing experiments by information gain and decision value, including when biased simulations are worth less than direct measurements.PDF
06Search, verifiers, and test-time computeScientific generation as search: comparing retraining, conditioning, ranking and intermediate steering while separating scores from scientific validity.Scientific generation as search: comparing retraining, conditioning, ranking and intermediate steering while separating scores from scientific validity.PDF
07Latent and continuous reasoningReasoning in latent or continuous representations and the relationship between internal computation and scientific decisions.Reasoning in latent or continuous representations and the relationship between internal computation and scientific decisions.PDF
08Recurrent, diffusion, and hybrid reasoningRecurrent, diffusion and hybrid reasoning, including particle steering, reweighting and the need for correction and independent evaluation.Recurrent, diffusion and hybrid reasoning, including particle steering, reweighting and the need for correction and independent evaluation.PDF
09Hierarchical reasoning and abstractionBuilding and reusing abstractions while preserving the assumptions that make a reasoning step valid.Building and reusing abstractions while preserving the assumptions that make a reasoning step valid.PDF
10Hierarchical planning and subgoal discoveryDecomposing investigations into subgoals, with a distinction between checking a plan and validating its scientific interpretation.Decomposing investigations into subgoals, with a distinction between checking a plan and validating its scientific interpretation.PDF
11Learning reasoning strategies with RLLearning reasoning strategies through reinforcement learning, including what the reward encourages and what it leaves untested.Learning reasoning strategies through reinforcement learning, including what the reward encourages and what it leaves untested.PDF
12Causal, mechanistic, and counterfactual reasoningConstructing causal alternatives and separating treatment effects, mechanisms and transfer through explicit interventions and assumptions.Constructing causal alternatives and separating treatment effects, mechanisms and transfer through explicit interventions and assumptions.PDF
13Scientific evidence and tool useUsing literature, databases and tools without confusing successful execution or a supporting citation with sufficient evidence.Using literature, databases and tools without confusing successful execution or a supporting citation with sufficient evidence.PDF
14Long-horizon research, memory, and self-improvementMemory and self-improvement across long research trajectories, and how persistent changes affect evaluation.Memory and self-improvement across long research trajectories, and how persistent changes affect evaluation.PDF
15Multi-agent scientific reasoningWhen collaboration between agents adds value beyond a single system with matched information and resources.When collaboration between agents adds value beyond a single system with matched information and resources.PDF
16Biomedical discovery systemsBiomedical discovery workflows that distinguish simulator predictions from measurements, and treatment effects from mechanisms and transfer.Biomedical discovery workflows that distinguish simulator predictions from measurements, and treatment effects from mechanisms and transfer.PDF
17Closed-loop experimentation and hypothesis revisionRevising competing hypotheses through noisy experimental results, with explicit likelihood updates, dependence checks and model-expansion limits.Revising competing hypotheses through noisy experimental results, with explicit likelihood updates, dependence checks and model-expansion limits.PDF
18Can AI Actually Discover? Novelty, Evidence and VerificationEvaluating whether AI systems make scientific discoveries: novelty, evidence and verification, with model accuracy, research behavior and scientific validity assessed separately.Evaluating whether AI systems make scientific discoveries: novelty, evidence and verification, with model accuracy, research behavior and scientific validity assessed separately.PDF
Opening slide from The Nature and Scaling of Scientific DiscoveryThe Nature and Scaling of Scientific Discovery10 lectures 2026A 10-lecture investigation of discovery as a computational process: hypothesis spaces, identifiability, experimental feedback, scaling laws and resource trade-offs. Examples from mathematical construction, program synthesis and molecular design examine what a predictive theory of scientific progress would need to explain.
The Nature and Scaling of Scientific Discovery: lecture topics and slides
No.LectureFocusSlides
01What Is a Scientific Discovery?What distinguishes discovery from solving a fixed task, and how interventions separate prediction, causal effects and competing mechanisms.What distinguishes discovery from solving a fixed task, and how interventions separate prediction, causal effects and competing mechanisms.PDF
02What Makes a Discovery Problem Complex?Separating computational difficulty, missing information and identifiability: when more computation cannot resolve a question.Separating computational difficulty, missing information and identifiability: when more computation cannot resolve a question.PDF
03Representations and the Geometry of DiscoveryHow representations reshape discovery, with symmetry motivating alternatives based on augmentation, invariant features and equivariant models.How representations reshape discovery, with symmetry motivating alternatives based on augmentation, invariant features and equivariant models.PDF
04Search, Experiments, and the Value of FeedbackTurning verbal hypotheses into explicit models, choosing experiments that distinguish them and updating beliefs from noisy evidence.Turning verbal hypotheses into explicit models, choosing experiments that distinguish them and updating beliefs from noisy evidence.PDF
05Sequential and Parallel DiscoveryWork, depth and adaptive dependencies: which parts of discovery can run in parallel and which must wait for evidence.Work, depth and adaptive dependencies: which parts of discovery can run in parallel and which must wait for evidence.PDF
06Scaling Laws and Resource Trade-offsTrading off data, computation and experiments while distinguishing predictive loss, success probability and validated discovery yield.Trading off data, computation and experiments while distinguishing predictive loss, success probability and validated discovery yield.PDF
07When Are Discovery Problems Equivalent or Transferable?When insights transfer between problems, and the different assumptions behind reductions, generalization and causal transport.When insights transfer between problems, and the different assumptions behind reductions, generalization and causal transport.PDF
08Proofs, Programs, and Molecules: Comparing Discovery Across DomainsComparing proofs, programs and molecules, including how Coscientist's tool-error repair differs from chemical measurement and scientific discovery.Comparing proofs, programs and molecules, including how Coscientist's tool-error repair differs from chemical measurement and scientific discovery.PDF
09Cumulative Discovery and Moving FrontiersHow discoveries accumulate and move the frontier, and why repeated claims or shared errors do not amount to independent corroboration.How discoveries accumulate and move the frontier, and why repeated claims or shared errors do not amount to independent corroboration.PDF
10Toward a Predictive Theory of Scientific ProgressWhat a predictive theory of scientific progress could claim, how to test it and where its predictions must remain conditional.What a predictive theory of scientific progress could claim, how to test it and where its predictions must remain conditional.PDF

Education

University of California, Berkeley

Masters in Data Science

Top-4 CS university in the US

Specialisation

NLP, AI agents, biomedical AI

National Technical University of Athens

Bachelors/Masters in Computer Science

Top CS university in Greece

Specialisation

AI, robotics, mathematics

Distinctions and awards

  • UC Berkeley honorary mention

    Perfect scores in selected UC Berkeley courses.

  • UC Berkeley merit scholarship

    Entrance scholarship based on prior academic performance.

  • Columbia University merit scholarship

    Offered an entrance scholarship based on prior academic performance (chose UC Berkeley over Columbia).

  • NTUA Kritikos Award

    Top academic performance in mathematics courses.

  • NTUA Papakyriakopoulos Award

    Top academic performance in mathematics courses.

  • Eurobank award

    Top 0.1% (top 40 among 100,000 candidates) in Greece's national university admissions exams.

  • Greek national mathematical olympiad reserve squad

    Selected for the national reserve squad.