Davies Meyer – home
    AI3 min read

    Model Evals (Model Evaluations)

    Model evals systematically assess AI models using defined tasks and criteria. They show performance under the tested conditions. A complete application also requires checks of data access, tools, interaction and actual outcomes.

    Model Evals (Model Evaluations) explained

    The key question is what must work well for this task. Correct specifications and clear language in product copy differ from the criteria for routing enquiries. A general does not replace that decision.

    Build suitable cases: ordinary tasks, missing information and situations requiring a stop or handover. Separate examples used for improvement from cases used for subsequent assessment. Otherwise, you may optimise only for familiar tasks.

    Automated checks, expert judgement and user tests answer different questions. Another language model can help grade results but needs assessable criteria itself. Record the model, data, instructions and grading method so comparisons remain meaningful.

    Examine serious failures and differences between relevant groups as well as averages. A change may improve one behaviour while worsening another. A passed assessment applies to its scope; it does not guarantee every future input.

    Examples

    Hypothetical application

    A team compares two configurations for product descriptions using the same approved data. It assesses factual accuracy, clarity and invented additions. After selecting a configuration, it tests further products not previously used for adjustment to see whether the findings hold.

    Key Points

    • Derive criteria from the real task.
    • Separate development examples from assessment cases.
    • Critically assess model-based graders too.
    • Examine averages and serious individual failures separately.

    Practical application

    Define acceptable outcomes and relevant failures. Keep test cases and grading traceable. Compare changes under suitable conditions and, for agents, also assess the state actually achieved.

    Useful measures

    Task-specific quality

    Fulfilment of previously defined substantive criteria.

    Relevant failures

    Frequency and severity of failures important to the application.

    Comparison stability

    Traceable differences between versions and groups of cases.

    Common mistakes

    • Treating a general score as approval for any application.
    • Mixing assessment and training material without control.
    • Hiding rare but consequential failures in averages.

    Sources and context

    Frequently Asked Questions about Model Evals (Model Evaluations)

    No. It reflects performance on those tasks and conditions. Suitability for your purpose needs assessment using relevant cases.

    Loading related terms…

    All Terms