Skip to content

Roadmap

This collects everything that was marked as "roadmap" in the document — what will come later. It is a direction of development, not a promise of specific dates: the order and scope may change.

The main step: recognizing students by voice (Voice ID)

Today the system knows how much students talked, but not "who exactly." The teacher can already record a sample of their own voice in the admin panel — that is how the system tells their speech from the students' (see the chapter "What is ControlAI"). The next big step is to distinguish the students themselves by voice: for each student a short voice sample is recorded once (at the start the teacher helps with this; eventually the student does it themselves), and from then on the system understands which student was speaking.

What this unlocks: - personal reports for each student (who talked how much, how many times they answered); - per-student progress over time; - the foundation for a personal student profile (see below) and reports to parents.

This is our key groundwork for the future — something that is hard for competitors to replicate, because it grows on the data a center accumulates.

Parent channel

Right now there are no parents in the system. The plan is a separate channel for parents (a portal or an app) with a brief report about their child: attendance and a brief summary of the lessons. It only makes sense after Voice ID — otherwise the report would be anonymous.

Student profile

Today students do not log in to the system. The plan is a personal student profile that the student sees themselves (and, where appropriate, their parent too). What it will contain:

  • lesson participation — how much the student talked, how many times they answered (after recognition by voice);
  • attendance and its trend;
  • week-by-week progress and personal achievements;
  • a link to the reports to the parent about this child.

The student profile relies on Voice ID: without recognition by voice, participation cannot be tied to a specific child. That is why it comes as the next step after recognizing students, and it appears not in the center's general admin panel but in a separate, protected form (data about minors — with elevated privacy requirements).

New video signals

On top of attendance and late arrivals — separate validated detectors: - how many students get distracted by a phone; - how many students are active and how many are passive throughout the lesson; - head on the desk (signs of sleep, fatigue, or boredom); - how often students raise their hand.

Finer-grained speech analytics

  • language share over the course of a lesson (how much Uzbek, how much Russian, how much English — by segment);
  • interaction pace (how many "teacher↔student" exchanges per 10 minutes);
  • distribution of speech across students ("3 students talked 90% of the time");
  • classroom audio-quality monitoring (catches a "dying" microphone in advance);
  • signals from speech: rude language, praise for students, addressing by name, distractions onto off-topic matters.

Deeper into methodology

  • lesson stages (warm-up → presentation → drilling → independent practice) as a separate signal and a check of the timings;
  • coaching tips for the teacher, generated from the rubric;
  • weekly summaries with trends, trend warnings (a rating dropping several weeks in a row), comparison with the center average.

Convenience and accuracy

  • schedule import from Excel/Google Sheets (less manual entry during onboarding);
  • continuous improvement of recognition of Uzbek-Russian speech — based on the corrections and annotations that accumulate in the system (the recordings themselves are still deleted on the schedule from the chapter "Privacy and trust"); the longer the system runs, the more accurate it gets.

What we deliberately do NOT do

Some things we have deliberately deferred — for reasons of accuracy, cost, or ethics: - open-ended "ask anything about the video" ("ask anything about any moment") — unreliable, expensive, and unacceptable for children's privacy; - assessment of emotions/engagement from faces — ethically contentious, especially for children; - a public teacher "leaderboard" — demotivating and at odds with the "helper, not surveillance" approach; - "live" analysis — the product analyzes a lesson that has already happened, it does not stream it.


Key takeaways from this chapter: ahead lies recognizing students by voice (and, built on top of it, personal reports and a parent channel), new video signals, finer-grained speech and methodology analytics, conveniences, and growing accuracy. And a number of things we deliberately do not do — for reasons of reliability, cost, and ethics.

→ Next: 17. Glossary — a short dictionary of terms.