Skip to main content

Responsible by design: Evaluating & testing AI tools for accuracy and reliability

1:00 pm – 2:00 pm ET
Webinar
1 hour

Generative AI has the potential to increase access to justice for self-represented litigants, but only through responsible use. To do this, courts must thoroughly test AI tools for accuracy, tone, and outcomes and then refine models based on results — before and after the tool is launched. 

This TRI/NCSC AI Policy Consortium for Law & Courts session will help participants understand the unique needs of evaluating and testing AI tools and introduce ideas and tools for planning and executing necessary evaluations across all stages of public-facing AI tool development. 

Topics include why evaluation is key and what it means to test an AI tool. Panelists will also cover acceptable error rates based on use, audience, and risk of harm; the critical role of high-quality content; and various testing methods and strategies used in the field.

Following this session, attendees will be able to:

  • Explain why testing GenAI requires additional steps to ensure the tool is providing accurate information, appropriate detail and correct tone.
  • Evaluate the true cost of GenAI tools in terms of content development, technology development, testing, and staffing and contractor costs.
  • Determine acceptable error rates based on the tool's use, audience, and baseline human error rates and testing results.
  • Distinguish between different types of testing methods, including creating testing rubrics and use of models and describe how testing volume and strategy may differ pre- and post-launch.

Register today

Panelists:

  • Margaret Hagan, executive director, Legal Design Lab, Stanford Law School
  • Keith Porcaro, assistant clinical professor of law, Duke University School of Law
  • Angela Tripp, program officer for technology, Legal Services Corporation

For more information, email Keeley Daye.