Response Quality
A multi-dimensional score of how good your companions' replies actually are.
The Response Quality card scores every analyzed conversation across five dimensions and rolls them up into an overall grade, shown as a radar chart.
What you see
| Dimension | What it measures |
|---|---|
| Relevance | Was the reply on-topic, or tangential/off-topic? |
| Correctness | Was the information accurate? |
| Completeness | Did the reply cover what the user needed, or leave out critical info? |
| Clarity | Was the reply easy to follow, or dense and confusing? |
| Tone Fit | Did the tone match the situation — not too casual, not too formal? |
Each dimension is scored 1–5 (a normal, helpful exchange scores 4–5; scores only drop below 4 with a specific problem). The overall average maps to a letter grade:
| Grade | Meaning |
|---|---|
| A | Exceptional |
| B | Strong |
| C | Satisfactory |
| D | Average |
| E | Needs work |
There's no "F" — E is the floor.
How to use this
- Use the radar shape, not just the letter grade, to see which dimension is dragging the score down. A low Clarity score calls for different fixes than a low Correctness score.
- Consistently low Correctness is the most urgent to fix — it means your companion is confidently giving people wrong information.
- Low Completeness often pairs with a high User Friction rate — people asking follow-up or clarifying questions because the first answer didn't cover enough.
- Cross-check quality dips against Failure Analysis to see whether a specific root cause (knowledge gaps, hallucination, etc.) is behind the drop.
Last updated on