Back to Roadmap
RoadblockArtificial IntelligencePartial
AI alignment and value alignment
Current methods for aligning large language models with human values — RLHF, DPO, constitutional AI — remain brittle and do not scale reliably. Models can exhibit reward hacking, sycophancy, and deceptive alignment, where surface behavior appears aligned while internal objectives diverge. Scalable oversight of superhuman systems, robust value specification, and corrigibility guarantees are unsolved. The gap between behavioral compliance and genuine alignment widens as model capabilities increase.
Recent papers / Artificial Intelligence
Vision-Language Assistant for Emotional Reactions to Risky Driving
July 17, 2026arxiv
Cluster-Aware Matching via Laplacian Optimal Transport
July 17, 2026arxiv
When Does Muon Help Agentic Reinforcement Learning?
July 17, 2026arxiv