Benchmark Whitepaper
Tencent HY3 Free Benchmark Whitepaper
A model-specific whitepaper on Tencent HY3 Free benchmark results for TuitionGoWhere quizzes and exam papers.
This whitepaper reviews how Tencent HY3 Free performed in the TuitionGoWhere benchmark. The benchmark used a separate evaluator model to score generated learning materials against language fit, syllabus fit, answer quality, exam format, notation, difficulty, missing-image risk, and other practical checks.
The results should be read as an internal quality signal, not a formal education certification. A score below 8.0 is treated as a review warning and is hidden from the main subject pages under the current quality policy.
Score Profile
Content Coverage
| Content type | Reports | Average score |
|---|---|---|
| Exam papers | 1,052 | 8.8 |
| Quizzes | 1,055 | 9.0 |
Criterion Scores
| Criterion | Score | Interpretation |
|---|---|---|
| Language suitability | 9.8 | Language was generally clear and level-appropriate in the tested sample. |
| LaTeX and notation handling | 9.9 | Very strong handling of formulas, symbols, and technical notation where present. |
| Artefact cleanliness | 9.9 | The output was usually clean, with few stray symbols or visible generation artefacts. |
| Syllabus adherence | 9.2 | Most content stayed close to the intended subject and topic, with some notable mismatches. |
| Answer explanation quality | 8.5 | Answers were often useful, though some needed clearer method-level working. |
| Question-template adherence | 8.1 | The main structural weakness was drift from the expected quiz or examination template. |
| Difficulty fit | 7.6 | Difficulty was the weakest measured area, with some resources too basic or poorly matched to the level. |
| Exam-paper format | 8.6 | Many papers had a workable structure, but marks, sections, and paper conventions were not always consistent. |
Main Findings
- HY3 Free performed strongly on language suitability, notation handling, and clean output. This makes it a useful candidate for broad first-pass generation where readable material and technical formatting matter.
- The strongest subject-level results appeared in Primary 3 Mathematics, A-Level Chemistry H1, Secondary 3 Biology, Secondary 1 Geography, and Secondary 4 Pure Biology.
- The weakest results were concentrated in Primary 5 and Primary 6 Higher Tamil, Primary 6 Higher Malay, O-Level Literature, O-Level History, and some upper-secondary Mother Tongue subjects.
- The benchmark recorded 2,107 completed reports from 2,224 generated artifacts. A further 117 reports were not scored because of evaluator-side failures, so the published averages should be read as results for the completed sample rather than the entire generation run.
Subject and Level Fit
- Strong fit: Primary Mathematics, Secondary sciences, A-Level Chemistry H1, and subjects that depend on clear notation or structured explanations.
- Use with review: Higher Mother Tongue, Mother Tongue, Literature, and history resources, where language, text type, and assessment structure need close control.
- Use with review: full exam papers, especially when exact marks, timing, section structure, and question-template fidelity are important.
Recurring Comment Themes
The benchmark comments were reviewed for repeated patterns. These themes overlap, so they should not be added together as independent failure counts.
- Difficulty mismatch was the most consistent weakness. Some upper-primary resources used content closer to an earlier primary level, while some advanced papers introduced structures that did not fit the intended examination.
- Language control was usually strong overall, but several Chinese, Malay, and Tamil resources were mainly written in English or used the wrong level of target-language vocabulary.
- Format drift appeared in Literature, History, General Paper, and some Mother Tongue papers, including incorrect section types, mark allocations, or short-answer structures where longer responses were expected.
- Some outputs contained internal revision notes or self-corrections inside the paper, which made them look like drafts rather than finished student materials.
- Answer keys were generally present, but selected mathematics, science, and structured-response questions still needed more explicit step-by-step working.
- Missing-image flags were common in diagram and visual-context questions. These are primarily a follow-up for the image-generation and attachment stages, not automatically a failure of the text generation itself.
Prompting and Workflow Implications
TuitionGoWhere does not publish its original internal prompts. The points below are high-level improvement directions based on the benchmark findings.
- Add a stricter level-and-subject gate before generation, with explicit exclusions for topics that belong to another stream or examination type.
- Require a final target-language ratio check for Chinese, Malay, and Tamil resources, and reject outputs whose main content is in English.
- Provide a compact paper-format checklist covering sections, marks, duration, response type, and total-mark arithmetic.
- Separate drafting from final output so internal corrections, planning notes, and alternative versions cannot appear in the student-facing paper.
- Add a difficulty calibration pass using representative questions from the intended level, followed by a targeted rerun for low-scoring subjects.
- Keep image placeholders and image generation as a separate pipeline stage, while preserving the original description so visuals can be regenerated later.