Overview

Content

People

Advisors

Guides

Apply

Support Us

This page is still under construction.

Theoretical Agent Foundations Applied Agent Foundations AI Alignment from Neuroscience Improved Preference Optimization
Probability theory, decision theory, propositional logic, measure theory, theoretical computer science Probability theory, formal logic, reinforcement learning theory Software engineering, ML frameworks (PyTorch/TensorFlow), experiment design Neuroscience fundamentals (neuroanatomy, fMRI analysis), computational modeling
Build and analyze idealized agents with provable value embedding; test theoretical guarantees. Build and analyze idealized agents with provable value embedding; test theoretical guarantees. Implement and scale agent architectures; validate in practical environments to measure robustness. Translate brain-derived value signals into algorithmic objectives; test consistency with human behavior.
Formal MDP frameworks, reward-learning algorithms, safety proofs in toy domains. Formal MDP frameworks, reward-learning algorithms, safety proofs in toy domains. Scalable RL systems, large-scale simulators, distributed training pipelines. fMRI decoding of value judgments, neural network models of decision-making.
Integrating theory of logic uncertainty or computational uncertainty with UDT setting, robustness under value uncertainty • Extending proofs to high-dimensional state spaces• Bridging gap between theory and real-world noise • Ensuring consistent reproducibility at scale• Balancing exploration vs. safety in dynamic settings • Low signal-to-noise ratio in neurodata• Aligning neural correlates with computational objectives