In one sentence

A busy pilot can still fail. The question is whether it improved a decision and whether the institution can operate it safely.

What Project Ardi showed

Project Ardi is a useful caution about metric definitions. Browser identifiers did not equal students, demo traffic contaminated early totals, and tool outages complicated interpretation. The stable public record is a 44-day pilot, a company-reported reach of more than 2,000 students, and 903 tracked conversations. More granular research required identity reconciliation and explicit exclusion rules.

Use a five-level measurement ladder

Level one is instrumentation: can events be trusted, identities reconciled, test traffic excluded, and outages marked? Level two is reach: who encountered and used the system? Level three is interaction quality: did the student provide context, challenge an answer, revise a choice, or abandon? Level four is workflow: did the system use the right source and complete the right handoff? Level five is institutional outcome: did anything durable change?

Each level depends on the one below it. An outcome dashboard built on ambiguous identity and missing outage labels produces precision without validity. Define the unit of analysis—person, session, thread, plan, appointment, or case—before reporting a rate.

Balance value, safety, and operational load

Value measures might include preparedness, decision confidence, time to a valid next step, plan revision, or successful resource arrival. Safety measures include unsupported claims, source freshness, missed escalations, inappropriate referrals, and disparate failure rates. Operational measures include review time, content corrections, staff escalations, uptime, and cost per completed workflow.

Do not collapse these into one score. A pilot can be useful and unsafe, safe and unusable, or popular and operationally unsustainable. Leaders need to see the tradeoff rather than a single green number.

  • Reach: unique people with a documented identity method.
  • Quality: valid, understood, revisable decisions.
  • Safety: error, uncertainty, escalation, and subgroup review.
  • Operations: ownership, correction time, reliability, and cost.
  • Outcomes: only claims supported by the study design.

Match every claim to the design

A single-site pilot without a comparison group can describe behavior and surface failure modes. It cannot establish that the tool caused a change in GPA, persistence, or equity. Interviews can explain why an interaction felt useful; they do not estimate population effects. Administrative outcomes can strengthen the study, but only when definitions, consent, matching, and confounding are addressed.

Write the limitations before results arrive. Predefine exclusions, missing-data treatment, identity reconciliation, and the difference between exploratory and confirmatory analysis. Honest limits increase the future value of the work because the next study knows exactly what remains unanswered.

Decision framework

Five questions to take into the room.

  1. 01

    Define the unit of analysis and identity method.

  2. 02

    Mark demos, outages, test accounts, and missing data.

  3. 03

    Pair behavioral metrics with reviewed interaction samples.

  4. 04

    Report value, safety, and operations as separate dimensions.

  5. 05

    Make only the outcome claims the study design supports.

Questions this note answers

  • Is adoption enough to call an AI pilot successful?
  • Which AI advising metrics are misleading?
  • How should test traffic and repeat users be counted?
  • Can a short pilot prove an effect on retention or grades?

About the evidence

Project Ardi was one 44-day product pilot at CU Boulder with no comparison group. It measured conversations and product behavior, not academic outcomes. “2,000+” is Ardvarq's company-reported estimate of student users; 903 is the tracked conversation count. Read the research notes and study limits.