Series
Responsible AI and Evaluation
How to design safety benchmarks, detect harmful outputs, build runtime guardrails, and measure what actually matters in production AI systems.
- 1
Designing Safety Benchmarks for LLMs: What Makes an Eval Good
Most safety benchmarks are gameable, distribution-shifted, or measure the wrong thing. Here's what separates a rigorous safety evaluation from a checkbox.
2024-06-19
- 1
Designing Safety Benchmarks for LLMs: What Makes an Eval Good
Most safety benchmarks are gameable, distribution-shifted, or measure the wrong thing. Here's what separates a rigorous safety evaluation from a checkbox.
2024-06-19
- 2
Privacy Leakage in LLMs: PII, Memorization, and Code Generation Risks
LLMs memorize training data. Under the right prompts, they reproduce it. Here's how memorization works, how to measure it, and the specific privacy risks in code generation models.
2024-06-19
- 2
Privacy Leakage in LLMs: PII, Memorization, and Code Generation Risks
LLMs memorize training data. Under the right prompts, they reproduce it. Here's how memorization works, how to measure it, and the specific privacy risks in code generation models.
2024-06-19
- 3
Runtime Guardrails: Architecture Patterns for Production AI Safety
Training-time alignment is not enough. Production AI systems need runtime layers that detect, intercept, and respond to harmful inputs and outputs. Here's how to build them.
2024-06-19
- 3
Runtime Guardrails: Architecture Patterns for Production AI Safety
Training-time alignment is not enough. Production AI systems need runtime layers that detect, intercept, and respond to harmful inputs and outputs. Here's how to build them.
2024-06-19
- 4
Red-Teaming vs Automated Evals: Tradeoffs and When to Use Each
Human red-teaming finds attacks automated evals miss. Automated evals achieve scale humans can't. Here's how to combine them, and what each can and can't tell you.
2024-06-19
- 4
Red-Teaming vs Automated Evals: Tradeoffs and When to Use Each
Human red-teaming finds attacks automated evals miss. Automated evals achieve scale humans can't. Here's how to combine them, and what each can and can't tell you.
2024-06-19