Most AI evaluation is borrowed from software QA: test coverage, error rates, latency and uptime. Those measures are necessary, but they are not enough for an agent that acts with real decision authority in real workflows. An agent that affects customers, staff or operations is closer to a member of the workforce than a software component, and it should be evaluated that way.
This programme teaches participants to apply the AI Agent Evaluation Framework (AAEF), developed by Ian Loe, to appraise agents against defined expectations, on a regular cadence, with governance that matches the authority they hold.
The Five Evaluation Dimensions
Task Competence
Does the agent perform its assigned tasks accurately and completely, measured against its own task specification rather than general benchmarks?
Judgement Quality
How does it handle ambiguity and edge cases? Does it know when to escalate, ask for clarification or decline to act?
Boundary Adherence
A binary governance measure. Did the agent stay within its sanctioned authority scope? Any violation requires review, regardless of outcome.
Collaboration Effectiveness
The clarity of its outputs, the usefulness of its escalations, and whether it supports rather than disrupts human oversight.
Improvement Trajectory
Is performance improving, stable or degrading over time? This is the signal for retraining, reconfiguration or retirement.
What You Will Learn
- Run structured AAEF appraisals using the appraisal template and scoring method
- Assign performance profiles (Operational, Developing, Restricted, Under Review) and know what each one means for oversight
- Set up a 30-day, 90-day and ongoing appraisal cadence for agents in production
- Define task specifications and authority boundaries that make evaluation possible
- Decide, with evidence, when to widen an agent's authority, restrict it or suspend it
- Connect agent-level evaluation to your wider governance and reporting
THE BENEFIT
Agents earn autonomy on evidence, not hope. AAEF gives you a repeatable, defensible way to show regulators, auditors and your board that every agent in production is performing, staying in bounds and being actively managed.
Who Should Attend
- Owners and product managers of AI agents in production
- Risk, audit and compliance teams overseeing autonomous systems
- Operations leaders whose teams work alongside AI agents
- AI engineers and FDEs who need to prove an agent is ready for more autonomy
AAEF operates within the governance envelope set by FRAME. FRAME without AAEF lacks agent-level rigour, and AAEF without FRAME lacks organisational context, so the two are designed to be used together.
Put Your Agents Under Review
Delivered as an in-house workshop, using the agents your organisation already runs or plans to deploy.
ENQUIRE ABOUT THIS PROGRAMME →