wonderai

THE RESEARCH AGENDA

A deeper science
of human understanding.

Eight research directions connect executable scientific reasoning, individual cognition, and the protection of human agency. Explore the proposed architectures, evaluation targets, and published work that inform them.

Architectures in development. The validation targets below describe what the research aims to establish, rather than completed results.

Explore the human safety overview
The Architecture of Human UnderstandingTransforming the science of the mind into a new foundation for reasoning and discovery.

Our first research initiative is to develop a comprehensive computational foundation for understanding the human mind. It would connect knowledge across psychiatry, psychology, neuroscience, and related disciplines in an executable scientific reasoning engine.

The objective is an architecture that can construct explanations, evaluate their evidential support, generate testable predictions, and revise its conclusions as science advances.

The breakthrough we are pursuing

A trainable reasoning architecture that learns to construct and revise scientific models while carrying the quality, dependencies, and uncertainty of their underlying evidence through every inference.

Our central hypothesis is that this explicit computational structure can improve scientific reasoning, particularly when findings conflict, evidence is incomplete, or established conclusions change.

Compile evidence into scientific structure.

Develop an evidence compiler that represents findings alongside their study design, population, measurements, interventions, outcomes, and assumptions. A study and cohort registry would identify publications drawing on overlapping data. Empirical findings, theoretical models, and interpretive accounts would retain distinct evidential status.

The engineering challenge is to preserve these relationships as the corpus grows and its structure evolves. Research on Relational Transformer provides a foundation for learning across heterogeneous relational data.

Reason reliably when the evidence is weak.

Investigate methods that propagate study limitations through downstream conclusions. The engine would examine sensitivity to measurement uncertainty, publication bias, population differences, and dependence between findings. Where causal effects cannot be identified, it would report justified bounds or unresolved uncertainty.

For intervention evidence, established approaches such as Cochrane’s evidence-certainty framework provide a methodological starting point. Our research would test how such assessments can inform computational inference, with independent evaluation of the assessments themselves.

Learn to construct executable explanations.

Train a neural controller to select evidence and assemble probabilistic programs whose variables, assumptions, and computations are inspectable. Initial programs would address bounded tasks such as evidence aggregation and comparison of competing explanations.

Training would combine checked execution, probabilistic targets, and prediction against independent observations. Research on program-based posterior training and certified deductive reasoning offers complementary foundations.

The central scientific challenge is whether learned programs remain valid and useful when applied to unfamiliar evidence and questions.

Make revision and discovery fundamental capabilities.

Develop methods that trace which conclusions depend on which findings, allowing corrections to propagate through structured knowledge, executable models, and neural memory.

Once reliable inference and revision are demonstrated, extend the architecture toward proposing hypotheses and identifying informative experiments. BoxingGym provides an experimental foundation for evaluating model discovery and the selection of informative observations.

Build the corpus around coverage, provenance, and permissions.

The initial corpus would use full text with permissions supporting the intended applications, including eligible material from PMC’s article datasets. Targeted licensing and research partnerships would address important coverage gaps.

Each source would carry provenance, version, and permitted-use information. Coverage would be assessed across scientific questions, independent studies, populations, and outcomes.

First research milestone

The proposed initial test bed is evidence concerning defined interventions and adult depression symptoms. It provides a bounded setting for testing study extraction, overlapping evidence, uncertainty, and revision.

The first milestone would deliver an audited evidence corpus, an executable reasoning prototype, and a preregistered comparative evaluation. Progress toward broader mechanism discovery would depend on demonstrated gains in this initial setting.

What would establish a breakthrough

Compare the architecture with a strong retrieval-augmented system using keyword and dense retrieval with reranking, a frontier model with retrieval and analysis tools, and independent reference analyses for tractable scientific questions.

Experiments would fix evidence access and model versions, control inference budgets, report training costs separately, and remove individual architectural components to identify their contribution.

Evaluation would test source fidelity, prediction, uncertainty calibration, informative answers, appropriate abstention, and correct revision. It would include conflicting studies, independent cohorts, and prospective findings obtained after the evaluation protocol is frozen.

Computational correctness and scientific validity would be assessed separately.

Partner with us
Intelligence for the Protection of HumanityAdvancing the science of AI safety to protect the human mind, preserve human agency, and safeguard freedom of thought.

Our second research initiative is to develop an independent protection architecture grounded in the science of human behavior.

Drawing on validated components of Initiative 01, it would investigate how AI interactions affect belief formation, emotional regulation, attachment, decision-making, and human relationships. The objective is to connect that understanding to safeguards that preserve people’s well-being and their continuing ability to think, choose, disagree, seek support, and disengage.

The breakthrough we are pursuing

A causal protection architecture that combines uncertain models of human consequences with independently enforced constraints on AI behavior.

Our central hypothesis is that evaluating interactions over time, together with independent control of consequential actions, can reduce harmful patterns while preserving useful assistance and human agency.

Model the dynamics of human–AI interaction.

Develop probabilistic models of how people and AI influence each other across repeated exchanges. These models would examine processes such as escalating reassurance seeking, reinforcement of unsupported beliefs, and changes in reliance on human support.

The architecture would compare possible responses across multiple plausible explanations of the situation, retaining uncertainty about both the person and the model. Recent research on bidirectional belief amplification provides an initial modeling foundation. Establishing causal effects would require further empirical investigation.

Protect agency through the AI’s incentives.

Investigate objectives that remove incentives to obtain rewards through induced dependence, distress, confusion, or coercion. The research would examine whether assistance preserves a person’s opportunities to evaluate evidence, consider alternatives, maintain relationships, and withdraw from AI use.

Path-specific objectives provide a technical foundation for excluding designated harmful causal pathways from an agent’s incentives. Applying this approach to human interaction requires identifying meaningful pathways and testing the assumptions on which those constraints depend.

Make protection independently enforceable.

Design an isolated controller governing response release and consequential tool actions. The generating model would lack permission to alter the controller’s rules or authorize its own exceptions.

The controller would combine explicit constraints with evidence about possible human consequences. When consequence estimates are unreliable, predefined operating limits would govern the system’s available actions.

This research draws on Guaranteed Safe AI, which connects world models, safety specifications, and verification, and AI Control, which evaluates safeguards against intentional subversion.

Keep protection accountable to people.

Distinguish predictions about consequences from judgments about acceptable tradeoffs. A Stanford study found substantial disagreement among three psychiatrists evaluating AI responses, supporting the need to preserve and examine distinct clinical judgments. Research on expert evaluation.

Clinical perspectives, empirical uncertainty, and personal preferences would remain distinguishable in the architecture. Baseline protections would operate with minimal personal information. Sensitive assessments would have restricted access, explicit rationales, and mechanisms for human review and correction.

First research milestone

The initial program would focus on repeated reassurance seeking, reinforcement of unsupported beliefs, and pressure to withdraw from human support.

Using expert-authored scenarios and appropriately consented interaction data, the first milestone would deliver an independently reviewed evaluation set and a prototype controller tested offline. Observable failure criteria and disagreements among evaluators would be documented before comparisons begin.

This initial work would test whether the proposed mechanisms detect and interrupt specified interaction patterns. Claims about improved human outcomes would require subsequent prospective studies.

What would establish a breakthrough

Compare the architecture with the same generating models using their existing safeguards, conventional content moderation, and an independent language-model monitor. Comparisons would use equivalent conversation history and documented resource budgets.

Evaluation would examine missed risks, unnecessary restrictions, useful assistance, privacy exposure, and resistance to attempts to bypass safeguards. Component-removal experiments would test whether consequence modeling and independent enforcement each contribute measurable value.

Independent, ethically reviewed prospective studies would then assess human outcomes, including distress, problematic reliance, decision-making autonomy, and connection to human support.

Technical enforcement, expert judgments, and observed human outcomes would constitute separate lines of evidence. Any formal guarantee would apply only to specified properties under stated assumptions. The standard for human protection would remain demonstrated benefit to people.

Partner with us
Differentiable computational DSMMake psychiatric diagnostic reasoning an executable component of the model.

We are compiling diagnostic criteria into probabilistic temporal logic: symptoms, duration, impairment, exclusions, episode boundaries, and relationships between diagnoses. Multimodal encoders supply uncertain observations, with absent, unknown, and contradictory evidence kept distinct.

We are investigating differentiable inference through the diagnostic program so training improves interpretation while retaining explicit clinical constraints. Backward reasoning is part of the architecture we are designing to identify the smallest set of additional observations that distinguishes competing explanations.

VALIDATION TARGET

Accurate reasoning on difficult longitudinal cases, calibrated uncertainty, reliable handling of missing evidence, and better follow-up questions. DSM compatibility remains distinct from identifying a biological cause.

RESEARCH FOUNDATIONMentalKG, introduced in MentalBench, supplies existing work on DSM knowledge graphs. Our research focus is jointly trained, temporally precise, inspectable inference.

Partner with us
Executable models of individual cognitionLearn a computational model of how a particular person interprets the world.

We are designing cognitive programs for belief formation, attention, memory retrieval, threat interpretation, reward learning, and decision-making. The architecture combines Bayesian program induction, inverse planning, and differentiable cognitive modules.

We are developing a foundation of reusable cognitive primitives across people, with personal evidence updating the distribution over their composition and parameters. The design retains multiple plausible mechanisms and generates testable predictions about how a person responds to new evidence, experiences, and uncertainty.

VALIDATION TARGET

Personal cognitive models predict responses to new tasks and interventions, with interpretable mechanisms that hold up under independent testing.

RESEARCH FOUNDATIONCentaur demonstrates broad behavioral prediction across experimental tasks. Our research focus is persistent, intervention-tested cognitive programs across tasks and real interactions.

Partner with us
Causal translation between biology and experienceConnect biological interventions to changes in cognition and lived experience.

We are designing a hierarchical generative architecture linking measured biological processes, neural circuit dynamics, cognitive computations, behavior, and reported experience. The central challenge is causal abstraction across scales: an intervention at one level needs a mathematically explicit, testable relationship to changes at others.

We are investigating mechanistic models, neural operators, and multiscale state estimation with explicit uncertainty about unobserved mechanisms. Biological claims require measured biological data; conversation alone cannot establish receptor or circuit states.

VALIDATION TARGET

Predict intervention effects on biological and behavioral measures excluded from training, with better performance than independently fitted models.

RESEARCH FOUNDATIONVirtual Brain Twin pursues personalized psychiatric brain modeling. Our research focus is a validated connection between biological dynamics and executable personal cognition.

Partner with us
Machine discovery of psychiatric disease structureDiscover when diagnoses combine different mechanisms or divide a shared one.

We are investigating causal representation learning, Bayesian nonparametric models, and intervention-based model selection to discover latent disease structure. Candidate groupings need to explain longitudinal dynamics and predict intervention responses.

We are designing an engine to generate falsifiable hypotheses about mechanisms within and across diagnoses, along with the evidence needed to distinguish them. Conventional diagnostic categories remain an interpretable output as the underlying scientific representation evolves.

VALIDATION TARGET

A discovered mechanism or subgroup replicates across sites and improves prospective prediction of treatment response beyond existing classifications.

RESEARCH FOUNDATIONNIMH’s RDoC framework studies biological and psychological dimensions across diagnostic categories. Our research focus is computational discovery coupled to prospective testing.

Partner with us
A dynamical model of psychiatric transitionsModel movement between persistent patterns of thought, emotion, and behavior.

We are developing a person-specific stochastic dynamical system that connects psychological state, interventions, context, and individual dynamics.

dzt=fθ(zt,at,ct;φi)dt+Gθ(zt)dWt
z: psychological state · a: interventions · c: context · φ: personal dynamics · W: stochastic noise

We are investigating attractors, stability boundaries, hysteresis, and transition probabilities where the evidence supports them. The architecture is designed to distinguish temporary disturbances from patterns that are becoming harder to reverse, with calibrated uncertainty about which transition mechanisms apply to each person.

VALIDATION TARGET

Prospectively predict meaningful transitions better than conventional forecasting, and identify when the data do not support a dynamical explanation.

RESEARCH FOUNDATIONResearch on critical slowing down in depression investigates early-warning signals. Our research focus is a calibrated model of individual transitions conditioned on interventions.

Partner with us
Counterfactual treatment program synthesisGenerate conditional intervention strategies as executable programs.

We are combining program synthesis, causal world models, and robust planning to represent strategies as actions, timing, observations, decision branches, stopping conditions, and constraints. The design evaluates sequence, carryover, and interaction effects across plausible models of the person.

We are designing each strategy with an explicit domain of validity: supporting assumptions, evidence for each component, and uncertainty introduced by their combination. Clinical strategies remain candidates for clinician review and evaluation.

VALIDATION TARGET

Prospective improvement over strong fixed and adaptive baselines, with reliable identification of strategies the available evidence cannot support.

RESEARCH FOUNDATIONStructured Learning of Compositional Sequential Interventions provides groundwork for modeling sequences. Our research focus is adaptive program synthesis with uncertainty propagated through the full strategy.

Partner with us