Metrics and What They Mean¶
In What ControlAI Can Do we listed what ControlAI measures. Here — what each number actually means, how to tell if it's good or not, and why it matters.
Important from the start. Numbers are a signal, not a verdict. What counts as "good" depends on the type of class (lecture, seminar, practical, lab), year of study, field, and the group's language. The benchmarks below are a starting point, not a rigid standard — every university tunes them to its own context. More on this at the end of the chapter.
Whether the class happened at all¶
Before talking about quality, ControlAI answers a simpler question — was the class held. For a university, this is the baseline level of oversight that used to require walking the halls and checking in person.
Class held / not held¶
- What it is. From the room's recording, the system determines whether a class took place: if no speech is heard in the room at the scheduled time, the class is flagged as not held — and this is visible to the rectorate, the dean, and the quality-control office (ichki ta'lim sifati nazorati bo'limi) without a single inspector in the hallway.
- Why it matters. Cancelled or "on-paper-only" classes are a direct loss of teaching hours and the biggest blind spot in traditional oversight. Here, every class in every room is recorded objectively and automatically.
Timing: when the class starts and ends¶
- What it is. How closely the class started and ended on schedule: the instructor was 12 minutes late, let the group go 20 minutes early — all of this is visible from the recording.
- Benchmark. It's a good sign when over 90% of classes start on time.
- Why it matters. A standard class is 80 minutes (2×40). Late starts and early finishes mean lost academic hours and disrespect toward students; over a semester, these "small" losses add up to weeks of lost instructional time.
Speech: what's audible in the class¶
Talk-time balance (students / instructor)¶
- What it is. What share of the time students spoke versus the instructor.
- Benchmark. The norm depends on the type of class. In a seminar or practical (amaliy) class, it's good when students talk 40–55% of the time: if the instructor takes up more than 70%, the practical session has likely turned into a monologue. In a lecture (ma'ruza), a high share of instructor speech is normal, not a warning sign.
- Why it matters. Material sticks when a student talks themselves — answers, reasons through a problem, defends a solution — not just listens. A skew toward the instructor in seminars and practicals is a common cause of weak mastery of the subject.
- Context. The type of class is known from the schedule, so lectures are correctly compared with other lectures, and practicals with other practicals. Today the same set of metrics applies to every class type, and whoever reads the report makes the adjustment for type; automatic calibration of benchmarks by class type is on the roadmap.
Instructor questions¶
- What it is. How many questions the instructor asked the room during the class.
- Benchmark. Around 12–18 questions in an 80-minute seminar or practical class is a healthy level of engagement. Fewer than ~8 and the class was probably run in "lecture and leave" mode. Lectures naturally have fewer questions, but zero questions is still a signal — even a lecture benefits from comprehension checks.
- Why it matters. Questions engage students, check understanding, and hold the attention of a large room.
Silence and pauses¶
- What it is. What share of the class passed in silence, and how long the longest pause was.
- Benchmark. Up to ~15% silence is normal. Short pauses are useful (students working through a problem, thinking about an answer), but long gaps — say, 6 minutes straight — signal that the class "stalled."
- Why it matters. A lot of silence means lost academic hours and, often, confusion or poor class organization.
Language of the class¶
- What it is. What language the class was conducted in, and in what proportions. A group's language of instruction — Uzbek, Russian, or mixed — is set when the group is created; speech recognition handles Uzbek–Russian code-switching in live speech with confidence. For foreign-language classes (language departments), it separately shows what share of the class was in the target language versus the students' native language.
- Benchmark. For language disciplines, the higher the share of the target language, the better the "immersion" — a typical quality bar is 70% or above. In earlier years of study, the share of the native language is naturally higher.
- Why it matters. How well the class language matches the group's language is part of teaching quality; for language departments, immersion is the core teaching tool.
Number of distinct speakers¶
- What it is. An estimate of how many different voices were heard during the class.
- Why it matters. An indirect check on engagement: if only two or three students answered during an entire seminar, the system will show it — even if the time-based balance looks fine.
Video: who was in the room (phase 2)¶
Being upfront about timing. The metrics in this section — attendance and lateness from video — belong to phase 2 of the system's development and aren't part of the current delivery. Everything above (whether the class was held, timing, speech metrics) already works today.
Attendance¶
- What it is. How many students were in the room versus how many were expected based on the group's roster.
- Why it matters. A drop in attendance is an early sign of trouble — with group discipline, with the schedule, or with the course itself; for the university, it's also a metric deans' offices already track.
Latecomers¶
- What it is. How many students arrived after the class started, and by how many minutes.
- Why it matters. Regular lateness signals a discipline issue, or something off with the group or schedule (for instance, a physically unrealistic transfer time between buildings).
Roll-up scores¶
On top of these metrics sit two "top-level" scores, each covered in its own chapter:
- Instructor rating — an overall performance score built from the metrics of their classes; how and why it's fair is covered in How the Instructor Rating Is Calculated. This is internal analytics for decisions and development, not a substitute for official certification.
- Methodology score for a class (1–10) — how closely a class matched the university's own observation form; it's calculated not from the metrics above but from checking the class against the form — see the "Methodology compliance" section in What ControlAI Can Do.
Portrait of a good class¶
Put together, here's what a healthy seminar / practical class looks like in numbers (benchmarks, not a standard; a lecture is read by its own norms, where a high share of instructor speech is expected):
| Metric | Good benchmark | Warning sign |
|---|---|---|
| Class held | every scheduled class | classes not held |
| Timing | on-time start and finish | regular lateness, letting the group go early |
| Student talk share | 40–55% | instructor > 70% |
| Instructor questions | 12–18 per class | fewer than ~8 |
| Silence | up to ~15% | long gaps, > 15% |
| Class language | matches the group's language (language disciplines ≈ 70%+ target language) | sharply diverges from the group's language |
| Attendance (phase 2) | stable / full | dropping class over class |
The benchmarks in the table are configurable by the university.
The key point: numbers are a signal, not a verdict¶
Three rules for reading the metrics:
- Context decides. A large lecture stream, a seminar, a practical, and a lab session all look different in the numbers. The same talk-time balance can be excellent in one case and poor in another — so compare classes of the same type. Automatic calibration of benchmarks by class type is on the Roadmap.
- A trend matters more than any single class. Everyone has an off class now and then. What's worth watching is the trend across the semester — three weak seminars in a row means a lot more than one.
- A number shows where to look — a person decides. ControlAI doesn't pass judgment. It highlights where attention is warranted; the head of department, the quality-control office, or the dean's office draws the conclusion, opening the recording or transcript if needed. For the instructor, those same numbers are personal stats and a development tool, not an inspection report.
The key takeaway from this chapter: every metric is a clear signal about whether a class happened and how it went, each with its own "good / warning" benchmarks. But they need to be read with the class type in mind and over time: ControlAI shows where to look; the decision always stays with a person.
→ Next: 06. How the Instructor Rating Is Calculated — what the overall score is built from, and why it isn't a "black box."