Skip to content

How the instructor rating is calculated

The instructor rating is the most sensitive number in the product. That's why the core principle here is transparency. The rating in ControlAI is not a subjective opinion and not a "black box": it is a clear formula built from measurable metrics — one that can be broken down to the last figure and explained in plain language.

A quick note on positioning, up front. The rating is the university's internal analytics — a tool for management decisions and instructor development. It does not replace formal certification (attestatsiya) and makes no claim to; at the university's discretion it can serve as supporting evidence for internal KPIs — but that is the university's choice, not a property of the system.

Don't confuse this with the lesson observation-form score (1–10) from the "What it can do" chapter — that one is about how a specific class session matches the university's own lesson observation form. Here we're talking about the instructor rating — an overall assessment of their work (on a 1–10 scale) built from the metrics of their class sessions.

What the rating is made of

The rating is a weighted sum of four factors. Each factor is drawn from the class-session metrics (see the Metrics chapter) and carries its own weight:

Factor Weight What it reflects
Speech balance 35% how much students speak in class versus the instructor
Questions 25% how actively the instructor engages students with questions
Language-of-instruction share 20% what share of the session runs in the group's language of instruction
Punctuality 20% whether sessions start and end on time

Why weighted this way: speech balance carries the most weight (35%) because it is the strongest signal of whether students are actually working in class or just listening. The weights are configurable: the university (the quality-control office or a department) can adjust them to its own requirements. The weights above are the current example; separate benchmarks for different session types (lecture, seminar) are covered in the roadmap.

The rating can always be "unpacked"

This is exactly what sets it apart from a "black box": the rating never arrives as one mysterious number. It can be expanded to see exactly what it's built from. Each factor is shown with:

  • a value (0–100),
  • its contribution to the change in the rating,
  • a plain-language explanation,
  • a color — what's good, what's dragging it down.

Example breakdown (an instructor whose rating is declining):

Factor Value Contribution Explanation
Speech balance 28 / 100 ▼ −0.42 The instructor started talking more (71% → 89%). The main cause.
Questions 30 / 100 ▼ −0.28 Average questions per session dropped from 19 to 8.
Language-of-instruction share 84 / 100 — 0.0 Stable, no issues.
Punctuality 94 / 100 — 0.0 94% of sessions started and ended on time.

A breakdown like this immediately shows not just that the rating dropped, but why — the instructor slipped into a monologue and stopped asking questions. And, most importantly, what to do about it: bring back the questions and let students talk. This turns the score from a label into a prompt for growth. This is exactly the kind of breakdown a head of department or mentor brings to a junior instructor — not "your class was bad," but "here are two specific numbers, and here's what to do about them."

Skills radar

The same data can also be viewed as a radar — a five-axis "star" of strengths and weaknesses. Each axis is calculated from the metrics by a clear rule:

Axis How the axis is calculated
Questions full score — around 28 questions per session
Speech balance full score — students speak roughly 55% of the time
Language-of-instruction share the higher the group's language-of-instruction share, the higher the axis
Engagement the less silence in the session, the higher the axis
Punctuality the more sessions on time, the higher the axis

Two clarifications so nothing gets confusing. The radar is slightly wider than the rating: a fifth axis — "Engagement" (silence) — is added to the four factors; it shows up on the radar but doesn't feed into the rating formula. And a "full score" on an axis is the level of an outstanding class, not the norm: the "healthy" 12–18 questions from the Metrics chapter already earns a solid mid-range score on the "Questions" axis.

One glance at the radar shows where an instructor is strong and what's worth working on.

A trend, not a one-off check

The rating is tracked week by week (typically over 8 weeks), so what shows up is a trend across the semester, not a random spike — unlike a one-off class visit by a review committee, which catches only one session out of dozens. Any change is explainable: the system shows which factor moved the rating, and by how much, over the period. "The rating has been slipping for three weeks straight because the instructor's share of talking time has grown" is a specific, verifiable conclusion — not a reviewer's impression. Semester-end reports — semester summaries by instructor, department, and faculty with semester-over-semester comparison — are covered in the Roadmap.

Why this is fair

  • The same for everyone. The same calculation applies to every instructor in a department — no favorites, and no dependence on who happens to be reviewing.
  • Built from facts, not opinion. Every factor is a measurable metric from actual class sessions held over the whole period, not an impression from one open lesson.
  • Fully transparent. The instructor sees exactly what the score is built from and what to improve — nothing hidden.
  • Configurable per university. Weights and benchmarks are tuned to the university's and department's requirements, not imposed from outside.
  • Accounts for context. The type of session (lecture, seminar, practical, lab) is known from the schedule, so a lecture is compared against other lectures, not against seminars: in a lecture the instructor naturally talks more. Automatically adjusting benchmarks by session type is on the roadmap; for now, whoever reads the report makes that adjustment themselves.

Rating, KPIs, and rectorate decisions

Within a department, the rating gives the head of department an honest picture: who is consistently strong, who needs support, who should be paired with a mentor. At the university's discretion, the rating can be linked to internal KPIs and incentive pay: because the rating is transparent, the incentive is transparent too — the instructor sees exactly which improvements in their sessions will raise both the score and the bonus. More detail is in the What it can do chapter and in Inside the admin panel. That said, the rating does not replace formal certification — it simply gives the university objective material for it, if it chooses to use the rating that way.

Important: the rating is for growth, not a label

  • The purpose of the rating is to help the instructor get better, not to punish them. It ties into mentoring junior instructors, goals, and actionable tips — not just a "grade from above."
  • There is no public leaderboard. The rating is a working tool for the head of department, the quality office, and the instructor themselves — not an excuse for a public competition that would demotivate part of the staff.
  • Any disputed number can be checked by opening the session transcript (and, within the first ~7 days, even the recording itself) — the final word always rests with a human.

Key takeaway from this chapter: the instructor rating is a transparent, weighted sum of four measurable factors (speech balance, questions, language of instruction, punctuality) that can always be unpacked down to a specific cause and a clear "what to improve." It is an internal tool for development and management decisions that does not replace formal certification — and the final word always rests with a human.

→ Next: 07. Reliability: no class session gets lost — what happens when a camera goes offline, a session isn't recorded, or a failure occurs.