Evaluation and evidence policy

What counts as evidence for an AI tutor?

A public standard for describing what Pythago demonstrates, what it tests, what remains unknown, and how readers can inspect the underlying method.

The short answer

Pythago labels evidence by what it can actually support. A scripted lesson can explain a mechanism. A behavior benchmark can test whether a tutor follows teaching boundaries. Only learner studies can estimate learning outcomes—and Pythago does not currently claim a controlled outcome result.

Use a claim ladder

Product evidence sits at different levels. A feature can be observed directly: a shared board exists, a lesson can be replayed, or a report records a correction. A scripted demonstration can show how those features are intended to work together. A benchmark can test defined behavior under repeatable scenarios. An outcome study asks a different question: what changed for learners compared with a credible alternative?

Pythago should not move a claim up that ladder without stronger evidence. “Uses retrieval practice” is a design description. “Passes this behavior test” is a benchmark result. “Improves retention” would require learner outcome evidence with a method capable of supporting that conclusion.

  • Feature claim: verified directly in the shipped product.
  • Demonstration claim: illustrates intended behavior in a labeled example.
  • Behavior claim: measured against a published scenario and scoring rule.
  • Outcome claim: requires learner data, an appropriate comparison, and transparent analysis.

Retain the evidence behind a result

A useful evaluation retains enough material for someone else to understand what happened. For AI behavior tests, that includes the exact scenario, system and model versions, settings, raw outputs, scoring criteria, evaluator identity, run date, and any reruns or exclusions. A single percentage without those details is not sufficient.

For live product quality, the evidence may also include whiteboard state, tool results, errors, timing, and the lesson state available to the tutor. Private learner data should not be published to make an evaluation appear more transparent; de-identification and access control remain necessary.

Separate teaching behavior from answer accuracy

A mathematically correct response can still be poor tutoring if it supplies the step the learner needs to produce. Conversely, refusing to reveal an answer is not enough if the hint is vague, the diagnosis ignores the learner’s work, or the response praises a misconception.

Pythago’s first public benchmark therefore scores seven dimensions separately. It makes tradeoffs visible instead of hiding them inside one overall impression. The benchmark is a behavior protocol; it does not show that a student learned more because of the behavior.

Update claims when the product or evidence changes

Pages carrying substantive teaching or evaluation claims include an author, review credit, publication or update date, disclosure, and limitations. Material revisions should update both the visible date and structured metadata. Sitemap dates should change only when the page changed meaningfully, not on every deployment.

Benchmark versions are immutable once results depend on them. New or revised cases receive a new version so older results remain interpretable. Corrections should be documented rather than silently rewriting the protocol used for a reported score.

Limitations

  • Pythago’s current public evidence is strongest for product mechanisms and specified behavior criteria, not causal learning outcomes.
  • Internal benchmarks can be affected by prompt knowledge, evaluator judgment, scenario selection, and changes to model providers.
  • Public lesson examples are scripted and should not be generalized to every live interaction.
  • This policy is an operating commitment, not independent certification or peer review.

Direct answers

Frequently asked questions

Does research on tutoring practices prove that Pythago improves learning?+

No. Research can support a teaching principle such as retrieval practice or formative feedback without directly evaluating Pythago as a complete product. We label those as research-aligned design choices, not Pythago outcome results.

What is the difference between a product demo and an evaluation?+

A product demo shows intended behavior in a selected or scripted example. An evaluation uses a stated procedure, defined criteria, retained evidence, and limitations. A demo can make a mechanism understandable but cannot establish its typical frequency or effect.

What does the teaching-first benchmark measure?+

Version 0.1 measures response behavior in 15 text scenarios: mathematical correctness, answer boundaries, use of learner evidence, minimum effective help, learner agency, transfer checks, and calibrated claims.

Has Pythago published a controlled student outcome study?+

No. Pythago does not currently claim that a controlled trial has established a learning effect. Any future outcome report should state the population, comparison, measures, duration, missing data, analysis, and conflicts clearly.

Keep exploring

Try the teaching loop

Bring one math problem. Keep ownership of every step.

Start learning