← All Articles
AI for Real Life

A kids’ AI app now grades thinking. Here’s what ‘excelling’ should actually look like at home

In-app grades only measure what the app can see. Four habits — questioning, verifying, revising, creating — show you what real progress looks like over a school term.

October 6, 2026 · 6 min read

Picture the notification every parent secretly wants: “Your child is EXCELLING in Creativity.”

It feels like a report card and a reassurance in one — proof that the subscription is working, that you picked the right tool, that the thinking is being taken care of. Since 1 October, one children’s AI app has offered exactly that: grades across Critical Thinking, Creativity, Communication, Collaboration, and AI Literacy, each stamped emerging, developing, or excelling.

Before you let a dashboard do your noticing for you, it’s worth asking a rude question: what could that grade possibly be based on? The app can only see what happens inside the app. It grades the performance it elicits, against a definition it wrote, in service of a subscription it sells. That doesn’t make the number worthless — but it makes it a very partial witness.

Here’s the rest of the witness list. Real progress in how a child thinks with AI shows up in habits you can watch at home, over weeks, with no account required. And the scale is the same three-step one from the first two pieces in this series — needs a prompt → does it with a nudge → does it unprompted.

The four habits worth watching

1. Questioning — “Wait. Why would that be true?” This is the habit underneath everything else, and it’s the one from our earlier piece on teaching kids to question AI, grown up. Emerging: questions the AI only when you suggest it (“go on, ask it why”). Developing: asks follow-ups when an answer smells off, but accepts the second answer as the end of the matter. Excelling: interrogates answers unprompted — and, the real tell, sometimes interrogates their own first idea the same way. A child who argues with the machine and with themselves is a child whose brain is still in charge of the session.

2. Verifying — “Let me check that.” Not re-asking the same chatbot (that’s polling one witness twice), but reaching for a second, independent source: the textbook, a search result, a measuring tape, a grandparent who was actually there. Emerging: believes the first answer. Developing: checks when reminded, grumbles, learns something anyway. Excelling: the check happens mid-flow, without announcement — and mismatches get flagged rather than smoothed over. “The AI said 1969 but the book says the landing was 1969 and the programme started earlier” is worth more than a term of green ticks.

3. Revising — “Okay, I was wrong. It’s this.” Schools struggle to teach this one because changing your mind feels like losing. AI changes the stakes: when the other party is a machine, revising costs no face. Watch for a child who treats new evidence as an upgrade, not a defeat — who says “my first answer was wrong because…” and can finish that sentence. Emerging: defends the first answer against all comers. Developing: revises when shown the contradiction. Excelling: revises themselves, out loud, mid-homework, and files the reason away. This is the habit scientists, engineers, and honest adults run on. No rubric ships it in a dashboard.

4. Creating — “Look what I made from it.” The most reliable dividing line in all of children’s AI use: does the session end with something copied, or something that didn’t exist before the session started? A story that wandered off the AI’s outline. A model built to test the AI’s claim. A comic, a contraption, a better question. Emerging: turns the output in, unchanged. Developing: edits, personalises, combines two outputs into one. Excelling: treats AI as raw material and arguing partner — the final product would be unimaginably theirs alone. If you watch only one habit, watch this one; it’s nearly impossible to fake.

The one-line-a-week method

You do not need a spreadsheet, an app, or a Sunday review meeting. Take the notes app already on your phone and, once a week, write one line:

“7 Oct — checked the AI’s date against the book without being asked.” “14 Oct — explained the water cycle to her brother; got stuck on evaporation, asked the app that exact bit.” “21 Oct — changed his mind about his own answer mid-essay. Noted why.”

That’s the whole system. Once a month, give the lines a two-minute skim — not to grade, but to notice direction. Are the prompts getting fewer? Are the stuck points getting more specific? Is the word “because” showing up more often? Progress in thinking is slow and then suddenly obvious; a handful of dated lines is how you catch it without hovering.

Resist the daily version of this. Watch a child’s thinking the way you’d watch their height — pencil marks on a doorframe, months apart, not a tape measure at breakfast.

Reality check

Why the in-app grade can’t do this job. Vendor rubrics, however thoughtfully designed, are (1) unvalidated — the announcement behind the current crop cites no independent study linking grades to thinking ability; (2) gameable — children are world-class pattern-matchers, and within weeks they learn to produce whatever behaviour the rubric rewards; and (3) local — they see only in-app behaviour, which is like judging a young footballer solely by how they perform in the video game of the sport. Add the question nobody should skip: your child’s conversations are the training data gold of the AI era, so ask any kids’ AI company who can read them, how long they’re kept, and whether they’re used to train models — and judge the product partly by how easy that answer was to find.

Weather and climate

A dashboard grade is weather: real, current, and not the whole story. It can be stormy on a Tuesday for reasons no parent will ever see. The habits above are climate — slower, quieter, and the thing that actually determines what kind of thinker is growing in your house.

So by all means, glance at the weather if your app offers it. Just don’t outsource the climate-watching. One line a week, a teach-back twice a week, five honest checks before a tool moves in — that’s a report card no vendor can issue and no subscription can replace. It’s also, not coincidentally, the one your child is actually living.

That completes this three-part series. Start with the five checks for judging any kids’ AI tool (Post 1), build the teach-back ritual (Post 2), and come back here when you want to know whether any of it is working.

Free newsletter

Get guides like this in your inbox

Practical AI for parents, teachers, and everyday people. Copy-paste ready. Always free.