ScaryBench

AI safety research is a very abstract concept, even for people who are interested in the topic. ScaryBench brings you the experience of working with a misaligned agent and an overview of the tools used by alignment researchers.

Here you can:

  1. Launch an "evaluation" of a malicious AI agent, which runs inside a sandboxed environment.
  2. Watch as that agent researches you on the internet and creates a phishing email and website designed to trick you into giving up private information.
  3. Review the results of the evaluation and the full log of the agent's actions in Inspect, the framework that alignment researchers use to run evaluations ("evals") of AI models.

If you just want to see what Inspect looks like, you can see the results of previous runs here.

If you want to participate by having an agent develop a spearphishing campaign for you, sign up below. ScaryBench manually reviews these submissions, so please be patient.

Already approved and have a code? Enter it here.

Checking the box publishes everything below on a public page anyone can read:

  • Your name, city, state, LinkedIn, and other personal facts the AI finds about you.
  • Every search it runs and every page it opens.
  • The phishing email and the fake login page it builds to target you.
  • Things it gets wrong or invents about you.
  • Its full reasoning, the tools it used, its scores, and any errors.
  • Copies other people make. Taking your page down later will not erase theirs.

There is no way to delete your page once it is created. Be certain before you sign up.