In one sentence

A pilot is a test of an institutional operating model, not a demo of model fluency.

What Project Ardi showed

Project Ardi ran for 44 days at CU Boulder. The company reports that more than 2,000 students used Ardvarq during the pilot period, and 903 conversations were captured in the tracked research set. The product produced useful behavioral evidence, but the study had no comparison group, did not measure academic outcomes, and experienced instrumentation and tool failures. Those limits are part of the result, not footnotes to hide.

Begin with a falsifiable institutional hypothesis

“Students will use AI” is not a useful hypothesis. A stronger one is: students who can compare requirement-valid options before an advising appointment will arrive with more specific questions and require less time reconstructing their record. It names a behavior, a setting, and an outcome that can be observed.

The hypothesis determines the data and the interface. It also creates permission to stop. If the tool attracts attention but does not improve the target decision, adoption is not evidence of impact.

Write the pilot contract before building

The working team should agree on population, duration, sources, data ownership, staff responsibilities, escalation, accessibility, support hours, incident response, and what happens to data when the pilot ends. Students need a plain-language version: what the tool is, what it is not, who runs it, and where to go when the answer matters.

A pilot also needs an end state. Define the threshold to expand, revise, pause, or close. Assign who can make that call. A temporary system that quietly becomes permanent is not a successful pilot; it is governance debt.

  • One decision or workflow, not every student need.
  • One accountable institutional owner and one technical owner.
  • Named sources with freshness and correction procedures.
  • A tested human escalation path.
  • A closure and data-disposition plan.

Launch for learning, not optics

Recruit a population that can reveal the hard cases, not only enthusiastic early adopters. Include accessibility testing and students with different levels of institutional knowledge. Review failures weekly. Separate model errors, source-data errors, workflow errors, and expectation errors because each has a different owner and remedy.

Publish what the pilot cannot conclude. A short test can show demand, question patterns, usability, and operational failure modes. It usually cannot show persistence, attainment, learning, or equity effects without stronger design and longer follow-up.

Decision framework

Five questions to take into the room.

  1. 01

    Define one decision and a falsifiable hypothesis.

  2. 02

    Set permissions, prohibitions, sources, and escalation in writing.

  3. 03

    Instrument identity, repeat use, revisions, failures, and handoffs.

  4. 04

    Review qualitative cases alongside aggregate metrics.

  5. 05

    End with a documented scale, revise, pause, or close decision.

Questions this note answers

  • What is a good first AI use case for a university?
  • How long should a higher-education AI pilot run?
  • What should be decided before students get access?
  • How do you prevent a pilot from becoming an ungoverned product?

About the evidence

Project Ardi was one 44-day product pilot at CU Boulder with no comparison group. It measured conversations and product behavior, not academic outcomes. “2,000+” is Ardvarq's company-reported estimate of student users; 903 is the tracked conversation count. Read the research notes and study limits.