How would you describe them?
Model-agnostic benchmark, generation and LLM-judging harness, and interactive viewer for evaluating selective censorship in language models.