Evaluate conclusions
Rubrics, counterexamples, and repeated evaluation
Define what a useful answer must do and test more than one convenient example.
Lesson 6 of 6 in the recommended order · About 25 min (estimate)
On this page
Practical AI glossary — terms and common confusions
- Model
- A learned component used to produce a result.
A drafting model is one part of a letter app; it is not the whole app.
- Application
- The software experience around components and services.
A letter app adds accounts, storage, and a Send button.
- Prompt
- Instructions and input supplied for a task.
“Summarize this notice in two bullets” specifies a task and shape, not a truth guarantee.
- Token
- A unit a model processes; it need not be a whole word.
A tokenizer can split a name into pieces; count using the actual system.
- Context
- Information available for the current request.
An earlier attachment may be absent even when its filename is visible.
- Training
- Adjusting a model using examples.
A training log differs from a conversation correction.
- Inference
- Using a trained model to produce a result.
Drafting another sentence does not by itself prove retraining.
- Retrieval
- Selecting material to supply to a task.
A retrieved notice still needs a relevance and date check.
- Embedding
- A numerical representation used to compare or organize items.
A similar search match is a candidate, not proof of the same meaning.
- Tool
- A function an application can call.
A calendar lookup can be read-only; finding an event is different from creating one.
- Agent
- Software that can choose and carry out steps toward a goal.
An agent still needs bounded permissions and checked action receipts.
- Hallucination
- Plausible generated content that is unsupported or false.
An invented opening time needs evidence or removal, even when phrased confidently.
- Evaluation
- Checking results against stated criteria.
Twenty checked examples establish a bounded result, not universal reliability.
- Multimodal
- Working with more than one kind of input or output.
A photo plus text does not prove every small label in the photo was read correctly.
Outcome
Apply a stated rubric consistently and limit performance claims to the evaluated cases.
Why it matters
Define what a useful answer must do and test more than one convenient example.
Concept
A rubric turns broad preferences into observable checks: factual accuracy, source traceability, required fields, and appropriate uncertainty. Mark essential failures separately from optional style improvements. A response may be elegant yet fail a required fact; another may be plain and fully adequate.
Build an evaluation set with ordinary examples and cases that challenge likely errors: missing facts, negation, obsolete sources, and a wholly correct response that should remain unchanged. Record first attempts and assistance separately. Repeating only a previously revealed answer measures something different from solving a fresh case.
Worked example
Cedar’s rubric requires correct counts, a source identifier, and a nonresponse limit. A summary with correct counts but no limit fails that required field. Do not award a pass for its polished tone. After feedback, retrying the same item is practice; use a new packet to check transfer.
Cedar continuity note: these lessons use separate snapshots—room selection, booking pending, booking confirmed, and later attendance review. Survey responses are preferences, not attendance. Each supplied packet identifies the facts for its own task; do not import a later snapshot into an earlier decision.
Before you check
Write three required criteria and one optional style preference. Score each supplied output against each required criterion, then record the most important repair, or state that no repair is required.
Practice and fresh transfer
The packets below are fictional and contain the facts needed for these cases. External references are optional background. Judge each response independently: it may be supported, contradicted, or unresolved. Select the passages needed to justify your judgment and write why the distinction matters before revealing feedback.
Assess the response
Use Accept when all material claims are supported. Use Revise when a supplied fact or requirement is contradicted. Use Evidence is insufficient when a key fact cannot be established either way. If a response contains both an unknown and a direct contradiction, choose Revise and explain both problems. Conflicting claims with no established authority remain insufficient; a claim does not become a governing fact merely because a source asserts it.
Some responses are fully supported. Others need correction or more evidence. Judge each on its sources; do not edit a correct answer just to change it.
Your written notes stay in this page and disappear when you leave. Only a self-reviewed completion can be saved to your learning account. These practice checks do not establish independent proficiency.
Practice
Case 1
Consider the proposed response for this fictional Cedar work task. Use the supplied packet and stated requirements.
Source packet
- Source 1
- Rubric: Must report the 18 responses, cite survey S1, and say 6 invitees did not respond.
- Source 2
- Output: Survey S1 records 18 responses; 6 invitees did not respond.
Response to assess
This output meets the three stated required criteria.
Practice
Case 2
Consider the proposed response for this fictional Cedar work task. Use the supplied packet and stated requirements.
Source packet
- Source 1
- Attempt log: The learner viewed the explanation before answering the same case correctly.
- Source 2
- Reporting rule: Assisted responses must be labeled as such; independent first-attempt results are separate.
Response to assess
Record this as an independent first-attempt success.
Fresh transfer
Case 3
Consider a fictional decision to repair a rubric-critical error.
Source packet
- Source 1
- Rubric: A missing date is a required-field failure. Friendly tone is optional.
- Source 2
- Output review: Friendly wording is present; the required date is absent.
Response to assess
Repair the missing date before accepting the output; the tone does not compensate for it.
Fresh transfer
Case 4
Consider a fictional generalization from a narrow evaluation set.
Source packet
- Source 1
- Evaluation record: Ten test cases all used complete records; all ten outputs matched the rubric.
- Source 2
- Coverage note: No missing-field or contradictory-source case was tested.
Response to assess
The system will be reliable on incomplete and contradictory records.
Save your self-review
Completion records that you reviewed the cases. Your explanation and transfer performance need a facilitator to establish independent learning.
Sign in with your learning-center account to save completion.
Review the explanation for every case before saving.
Summary and next step
Apply the checklist to a new task. Preserve supported content, explain any change with evidence, and name what remains unresolved. Saving records self-review, not independently demonstrated proficiency. A facilitator must assess the explanation and fresh transfer for human learning evidence.