Let’s be honest: the phrase “future-ready skills” has been beaten to death by every consulting firm, TED talk, and education conference since about 2018. Critical thinking, collaboration, creative thinking — we all nod along, then go back to teaching to the test because that’s what actually gets measured.
But Google Research, in partnership with NYU, just dropped something that might break that cycle. It’s called Vantage, and it’s a research experiment now available on Google Labs that uses generative AI to assess exactly those squishy, hard-to-measure skills. And the kicker? The AI’s scoring is on par with human experts.
The problem with measuring what matters
Standardized tests are great for math facts and grammar rules. They’re terrible for figuring out if a student can navigate a conflict, build on someone else’s idea, or think on their feet when a plan falls apart. Real-world skills don’t happen in a vacuum with multiple-choice bubbles.
You can’t fairly assess conflict resolution if the group never disagrees. You can’t test creative collaboration if everyone just agrees with the first idea. Human-based assessment works — one-on-one interviews, role-playing, observed group work — but it’s expensive, inconsistent, and impossible to scale across thousands of students.
Vantage tries to solve this by putting learners into dynamic, multi-party conversations with AI avatars. Think of it as a sandbox where the AI plays the role of teammates, managers, or even difficult stakeholders. The setup is controlled enough to be standardized, but flexible enough to feel authentic.
How the AI runs the show
The architecture is clever. There’s an “Executive LLM” that watches the conversation unfold in real time. It has a rubric — the same kind a human evaluator would use — and it steers the AI avatars to create situations that let the learner demonstrate specific skills.
Need to see if someone can handle pushback? The Executive LLM tells an avatar to challenge an idea. Want to test creative synthesis? It introduces a constraint or a conflicting viewpoint. By the end of the conversation, the system has gathered enough data to score the learner across multiple dimensions, all without the learner feeling like they’re being tested.
This isn’t a chatbot that just asks questions. It’s an adaptive assessment engine that dynamically generates scenarios based on what’s already happened. That’s a step beyond anything I’ve seen in edtech so far.
The numbers check out
The study with NYU compared Vantage’s scores against human expert evaluations. The correlation was strong enough that the researchers are confident the AI is measuring the same constructs a human would. That doesn’t mean it’s perfect — no assessment is — but it’s better than the alternatives we have now.
I’d have liked to see more detail on where the AI struggled. Was it worse at scoring nuanced social dynamics? Did it over-index on verbosity or confidence? The tech report is available, so I’ll dig into it, but the initial results are promising.
The bigger picture
Vantage is still a research experiment. It’s not replacing SATs or college admissions tomorrow. But it points to a future where assessment isn’t a separate, stressful event — it’s embedded in the learning process itself. You practice a skill, and the system gives you feedback in real time.
That’s the dream, anyway. The reality is that schools move slowly, budgets are tight, and any new tool has to prove it’s worth the investment. Google’s track record with experimental products is mixed — some become core tools, others vanish into the Labs graveyard.
Still, this is one of the more thoughtful applications of generative AI I’ve seen in education. It’s not just generating essays or answering homework questions. It’s trying to do something genuinely hard: measure the things we say we value but never actually test.
If Vantage survives the transition from experiment to product, it could change how we think about assessment. If it doesn’t, at least it showed that the AI can do it. The question is whether we’re ready to let it.
Comments (0)
Login Log in to comment.
Be the first to comment!