Benchmark Whitepaper

Tencent HY3 Free Benchmark Whitepaper

A model-specific whitepaper on Tencent HY3 Free benchmark results for TuitionGoWhere quizzes and exam papers.

Singapore

This whitepaper reviews how Tencent HY3 Free performed in the TuitionGoWhere benchmark. The benchmark used a separate evaluator model to score generated learning materials against language fit, syllabus fit, answer quality, exam format, notation, difficulty, missing-image risk, and other practical checks.

The results should be read as an internal quality signal, not a formal education certification. A score below 8.0 is treated as a review warning and is hidden from the main subject pages under the current quality policy.

Score Profile

Total reports 2,107
Quiz and paper reports 2,107
Overall score 8.9
Quiz and paper average 8.9
Below 8.0 254
Below-8 rate 12.1%

Content Coverage

Content type Reports Average score
Exam papers 1,052 8.8
Quizzes 1,055 9.0

Criterion Scores

Criterion Score Interpretation
Language suitability 9.8 Language was generally clear and level-appropriate in the tested sample.
LaTeX and notation handling 9.9 Very strong handling of formulas, symbols, and technical notation where present.
Artefact cleanliness 9.9 The output was usually clean, with few stray symbols or visible generation artefacts.
Syllabus adherence 9.2 Most content stayed close to the intended subject and topic, with some notable mismatches.
Answer explanation quality 8.5 Answers were often useful, though some needed clearer method-level working.
Question-template adherence 8.1 The main structural weakness was drift from the expected quiz or examination template.
Difficulty fit 7.6 Difficulty was the weakest measured area, with some resources too basic or poorly matched to the level.
Exam-paper format 8.6 Many papers had a workable structure, but marks, sections, and paper conventions were not always consistent.

Main Findings

  • HY3 Free performed strongly on language suitability, notation handling, and clean output. This makes it a useful candidate for broad first-pass generation where readable material and technical formatting matter.
  • The strongest subject-level results appeared in Primary 3 Mathematics, A-Level Chemistry H1, Secondary 3 Biology, Secondary 1 Geography, and Secondary 4 Pure Biology.
  • The weakest results were concentrated in Primary 5 and Primary 6 Higher Tamil, Primary 6 Higher Malay, O-Level Literature, O-Level History, and some upper-secondary Mother Tongue subjects.
  • The benchmark recorded 2,107 completed reports from 2,224 generated artifacts. A further 117 reports were not scored because of evaluator-side failures, so the published averages should be read as results for the completed sample rather than the entire generation run.

Subject and Level Fit

  • Strong fit: Primary Mathematics, Secondary sciences, A-Level Chemistry H1, and subjects that depend on clear notation or structured explanations.
  • Use with review: Higher Mother Tongue, Mother Tongue, Literature, and history resources, where language, text type, and assessment structure need close control.
  • Use with review: full exam papers, especially when exact marks, timing, section structure, and question-template fidelity are important.

Recurring Comment Themes

The benchmark comments were reviewed for repeated patterns. These themes overlap, so they should not be added together as independent failure counts.

  • Difficulty mismatch was the most consistent weakness. Some upper-primary resources used content closer to an earlier primary level, while some advanced papers introduced structures that did not fit the intended examination.
  • Language control was usually strong overall, but several Chinese, Malay, and Tamil resources were mainly written in English or used the wrong level of target-language vocabulary.
  • Format drift appeared in Literature, History, General Paper, and some Mother Tongue papers, including incorrect section types, mark allocations, or short-answer structures where longer responses were expected.
  • Some outputs contained internal revision notes or self-corrections inside the paper, which made them look like drafts rather than finished student materials.
  • Answer keys were generally present, but selected mathematics, science, and structured-response questions still needed more explicit step-by-step working.
  • Missing-image flags were common in diagram and visual-context questions. These are primarily a follow-up for the image-generation and attachment stages, not automatically a failure of the text generation itself.

Prompting and Workflow Implications

TuitionGoWhere does not publish its original internal prompts. The points below are high-level improvement directions based on the benchmark findings.

  • Add a stricter level-and-subject gate before generation, with explicit exclusions for topics that belong to another stream or examination type.
  • Require a final target-language ratio check for Chinese, Malay, and Tamil resources, and reject outputs whose main content is in English.
  • Provide a compact paper-format checklist covering sections, marks, duration, response type, and total-mark arithmetic.
  • Separate drafting from final output so internal corrections, planning notes, and alternative versions cannot appear in the student-facing paper.
  • Add a difficulty calibration pass using representative questions from the intended level, followed by a targeted rerun for low-scoring subjects.
  • Keep image placeholders and image generation as a separate pipeline stage, while preserving the original description so visuals can be regenerated later.

References