Solaris Research is a technical AI safety initiative by researchers at ETH Zürich. We build the science to detect, understand and prevent failures in frontier models: independent, open safety science the world can rely on.
Frontier models are deployed in new contexts at massive scale1, yet we have almost no science for how risk profiles shift as models undergo continuous narrow fine-tuning, or are put in harnesses, post-release. We study this gap with a concrete research bet: emergent misalignment such as deception, sycophancy and evaluation awareness leave latent low-dimensional traces inside models, allowing not only for mechanistic reversal but also scaling-law prediction from training dynamics.
In our recent oral position paper at ICML 2026 (top 0.7%), Anthropomorphic Misalignment Research Needs Stronger Evidence, we show that many scientific claims on safety are not supported by sufficient evidence, proposing explicit evidence levels for grading how strong a claim truly is. Our wider work spans theory and practice: theoretical results on LLM safety being low-dimensional, practical red-teaming on prompt injection, geometric polytope- and probe-based steering, reproducible analysis of reward-model optimisation and RLHF misleading humans, benchmarks for cryptographic schemes, and evaluation with differential pullback. We develop and release open-source software and run Apertus Claritas, the public hub for safety research on open-weight models.
From evidence to governance. Frontier safety is not only a technical problem for science to solve, nor a Western preference, but a universal right, and should be treated as such. Securing it calls for technical governance grounded in open, neutral and reproducible research; the kind of independent work we are on a mission to do.
