Bollwerk Labs
An applied defense lab in Zurich.
Transforming how companies and nations detect and eliminate threats. We deploy inside institutions, and what the agents learn there becomes research that ships back into the product.
Four unsolved problems stand between AI agents and real authority.
- Long-horizon investigation. Agents maintain and revise hypotheses across many tools, systems, and time periods. We build environments and benchmarks for investigations that unfold over hundreds of steps.
- Adversarial robustness. Evidence may be deceptive, generated, incomplete, or deliberately poisoned. We build synthetic adversarial cases and red-team environments.
- Policy-constrained action. Agents must respect permissions, regulations, and institutional risk boundaries. We build a policy engine that evaluates consequential actions at call time.
- Verifiable decisions. Consequential decisions must be reconstructable and challengeable. We build evidence lineage and replayable traces.
Agents need a bank to practice on. So we built one.
Frontier is a connected synthetic institution with customers, transactions, alerts, cases, analyst decisions, and security telemetry. We attack it with synthetic identities, credential replay, transaction laundering, and poisoned evidence.
We Built a Fake Bank ยท The Advisor Persona
Every decision makes the agents smarter.
Foundry turns closed cases and confirmed threats into signal, synthesizes unseen typologies, and promotes a candidate only after it beats the last agent across the institution's history. Synthetic cases stay inside the customer's environment.
Thousands of judged conversations before a real one.
Simulated environments combine personas, scenarios, mocked tools, and adversarial conversations. Each run measures the decision, evidence trail, policy compliance, and investigation path.
Built in Zurich at the intersection of frontier AI and Europe's regulated institutions.
A decade supervising Swiss banks shapes the controls. Production ML systems at Moderna and Novartis shape how models are evaluated and shipped. ETH recruiting relationships keep the work close to Zurich's research community.
Selected writing
- We Built a Fake Bank
- Adding a Second Audience to Frontier
- ExploitBench
- Benchmarking Red- and Green-Flag Extraction
- Squire Sandbox Agent