ZeroShotMind

Series

Responsible AI and Evaluation

How to design safety benchmarks, detect harmful outputs, build runtime guardrails, and measure what actually matters in production AI systems.

Defining Frontier
responsible-aisafetyevaluationbenchmarking
  1. 1

    Designing Safety Benchmarks for LLMs: What Makes an Eval Good

    Most safety benchmarks are gameable, distribution-shifted, or measure the wrong thing. Here's what separates a rigorous safety evaluation from a checkbox.

    2024-06-19

  2. 1

    Designing Safety Benchmarks for LLMs: What Makes an Eval Good

    Most safety benchmarks are gameable, distribution-shifted, or measure the wrong thing. Here's what separates a rigorous safety evaluation from a checkbox.

    2024-06-19

  3. 2

    Privacy Leakage in LLMs: PII, Memorization, and Code Generation Risks

    LLMs memorize training data. Under the right prompts, they reproduce it. Here's how memorization works, how to measure it, and the specific privacy risks in code generation models.

    2024-06-19

  4. 2

    Privacy Leakage in LLMs: PII, Memorization, and Code Generation Risks

    LLMs memorize training data. Under the right prompts, they reproduce it. Here's how memorization works, how to measure it, and the specific privacy risks in code generation models.

    2024-06-19

  5. 3

    Runtime Guardrails: Architecture Patterns for Production AI Safety

    Training-time alignment is not enough. Production AI systems need runtime layers that detect, intercept, and respond to harmful inputs and outputs. Here's how to build them.

    2024-06-19

  6. 3

    Runtime Guardrails: Architecture Patterns for Production AI Safety

    Training-time alignment is not enough. Production AI systems need runtime layers that detect, intercept, and respond to harmful inputs and outputs. Here's how to build them.

    2024-06-19

  7. 4

    Red-Teaming vs Automated Evals: Tradeoffs and When to Use Each

    Human red-teaming finds attacks automated evals miss. Automated evals achieve scale humans can't. Here's how to combine them, and what each can and can't tell you.

    2024-06-19

  8. 4

    Red-Teaming vs Automated Evals: Tradeoffs and When to Use Each

    Human red-teaming finds attacks automated evals miss. Automated evals achieve scale humans can't. Here's how to combine them, and what each can and can't tell you.

    2024-06-19