Figma now has a lot of AI: an agent, Code Layers, Motion, Weave, generative plugins, shaders, 3D transforms, and AI in FigJam and Slides. Every one launched with a confident demo, and demos are designed to flatter. The only way to know which of these features earn a place in real work is to run genuine work through all of them and score the results on the same rubric. This piece gives you that method: a scorecard you run on one real project, every feature, honest marks for what helped, what didn't, and what to actually use. Because these are recent features, the exact surface will keep evolving, but the rubric and method hold regardless.
If you want the single-tool comparison against other AI design tools, that's a different piece; this one stays inside Figma and asks which of its AI features are worth your time. It pairs naturally with the six-months-later adoption data on which features designers actually kept using.
What the scorecard will show
Figma now has a lot of AI: an agent, Code Layers, Motion, Weave, generative plugins, shaders, 3D transforms, and AI in FigJam and Slides.
The pattern you'll almost certainly find is that these features don't deserve one collective verdict. Some are genuinely useful in daily work, some are situationally useful, and some are impressive in a demo and rarely worth reaching for. Lumping "Figma AI" into a single thumbs-up or thumbs-down misses the point; the value is knowing which specific features to use for which specific jobs, which is what the per-feature scores give you.
The method and the rubric
A scorecard is only as trustworthy as the test behind it, so here's how to run it.
The method: pick one real project, ideally something with genuine complexity (real states, real data, more than a couple of screens), and put every Figma AI feature through it on the same project, so the features are judged against the same real work rather than cherry-picked demos. Using one consistent project is what makes the scores comparable, because each feature faces the same reality.
The rubric: score each feature on three dimensions, because "is it good" is too blunt to be useful.
Speed: does it actually save time versus doing the task without it? A feature that produces a result you then spend longer fixing than doing it yourself scores low here, however impressive the output.
Quality: is the output genuinely good, or plausible-but-flawed? This is where you apply real design QA, checking for the states, accessibility, system fidelity, and correctness that a demo glosses over, the same standard as design QA on any AI output.
Reliability: does it produce good results consistently, or is it a coin flip you can't depend on? A feature that's brilliant one time and useless the next is hard to build into a workflow, and that inconsistency matters as much as peak quality.
Score each feature on all three, with evidence, and the scorecard becomes a genuine guide rather than an opinion.
Feature by feature: what to test and what to expect
For each feature, here's what to put it through and what to expect, framed as informed prediction to confirm with your real testing, not as results. Score each on speed, quality, and reliability.
The Figma agent. Test it on generating and modifying real screens for your project, with proper briefs. Expect it to be strong at producing plausible first drafts fast and weaker at the particular and at edge cases, with quality tracking how well you prompt it. The agent is likely the most-used feature, so score it carefully. See the guide on prompting it well for how much input changes the result.
Code Layers. Test it by bringing real code into the file and converting a real component both ways. Expect strength on the presentation layer and weakness on integration, with value scaling with how clean your design-to-code connection is. Likely high ceiling, adoption-dependent.
Figma Motion. Test it by building real UI animations and microinteractions for your project. Expect this to score well, because it fits an existing workflow (designers already animate) and keeps the work in the file, low friction, clear utility.
Weave. Test it by building a repeatable workflow for a multi-step, rule-governed task in your project. Expect it to score well for genuinely repetitive production sequences and poorly if forced onto one-off or judgment-heavy work. Situational, high for the right task.
Generative plugins. Test by generating a small, specific plugin for a real annoyance in your workflow. Expect good results on simple, well-scoped tasks and diminishing returns on complex ones. Useful for the niche, personal tool.
Agent skills. Test by building a reusable skill for a recurring task and running it across the project. Expect value that compounds with reuse and depends heavily on how well you specify the skill. Score reliability especially, since a skill runs repeatedly.
Shaders. Test by adding effects to presentational surfaces in the project. Expect genuine value for subtle brand and background work and poor value (and performance cost) in functional UI. Narrow but real when used with restraint.
3D transforms. Test by adding depth where it might genuinely help (tactile microinteractions, a showcase) and where it doesn't (functional screens). Expect a narrow set of good uses and a wide set of tempting bad ones, plus performance and handoff costs. Likely low for most product work.
FigJam and Slides agents. Test the FigJam agent on setting up and cleaning up a real session, and the Slides agent on structuring and formatting a real deck. Expect usefulness for mechanical setup, cleanup, and formatting, and weakness at generating the actual thinking or specific substance.
The reason to test each on the same project and the same rubric is that it surfaces the real distinctions the demos hide: which features you'll reach for daily, which only for specific jobs, and which you can safely ignore.
What to actually use, and what to ignore
Once the scores are in, the synthesis is the payoff, and it holds in shape even before you have the exact numbers, because it follows from how these features fit real work.
Expect a tier of daily-use features: the ones that fit an existing workflow at low switching cost and produce reliable value, likely the agent (used well) and Figma Motion for most designers. These are worth learning properly, because you'll use them constantly.
Expect a tier of situational features: powerful for a specific job and irrelevant otherwise, likely Weave (for repetitive production), generative plugins (for niche tools), agent skills (for recurring tasks), Code Layers (for teams with a clean design-to-code setup), and shaders (for brand surfaces). Learn these when your work matches the job, and not before.
Expect a tier of rarely-worth-it features: impressive in a demo, narrow in practice, likely 3D transforms for most product work and, for many teams, the FigJam and Slides agents beyond mechanical help. Know they exist; reach for them only in the narrow cases they suit.
The meta-finding, which is the real value of testing everything against one project, is that "Figma AI" is not one thing to adopt or reject but a toolkit to be selective about, and the designers who get the most from it are the ones who know precisely which features to use for which jobs and which to skip, rather than either dismissing all of it or trying to use all of it. That selectivity is the same discipline that separates useful AI adoption from AI-for-its-own-sake across the board, the theme of designing AI features users trust.
Why demos mislead, and why testing beats them
It's worth being explicit about why a scorecard like this is necessary at all, rather than just reading the launch demos. A demo is built to show a feature at its best: a clean input, an ideal case, a task the feature was designed to nail, presented in a controlled setting. Real work is none of those things, messy data, edge cases, awkward requirements, the pressure of an actual deadline. The gap between demo performance and real performance is exactly what a demo is engineered to hide and what a scorecard is built to expose. A feature can look transformative in a thirty-second clip and score poorly on real work because the demo showed the one case it handles well. This is why testing on your own genuine project, not a demo-friendly one, is the whole point: it puts the feature in the conditions it'll actually face, which is the only test that predicts whether it'll help you.
How to score fairly
A scorecard is only trustworthy if the scoring is fair, so a few principles keep it honest. Compare each feature against the real alternative, not against nothing: a feature "saves time" only if it beats how you'd do the task otherwise, so a fast feature that produces work you spend longer fixing scores low on speed despite feeling quick. Judge quality with real QA, not a glance, because plausible-looking output that fails on states or accessibility isn't quality, and scoring it high because it looks finished defeats the purpose. Test reliability across several attempts, not one, since a feature that's brilliant once and useless twice can't be built into a workflow, and a single lucky result overstates it. And score against your actual context: a feature that's useless for your work isn't a bad feature universally, it's a poor fit for you, so note that distinction rather than marking it down absolutely. Scoring this way produces a card that predicts real value rather than one that rewards good demos.
Common testing mistakes to avoid
Three mistakes undermine a test like this, and naming them keeps your scorecard credible. The first is testing on demo-friendly work, cherry-picking the clean, ideal case that flatters the feature, which just reproduces the marketing; use genuinely messy real work instead. The second is judging on first impression, marking a feature high because the output looked impressive without running it through real QA, which is how plausible-but-flawed output gets overrated. The third is testing once and generalizing, drawing a confident conclusion from a single result rather than several, which overstates both the good and the bad. Avoid these three and your scores mean something; fall into them and you've produced a prettier version of the demos you were trying to see past.
What the scorecard changes about how you work
The payoff of doing this isn't just the article; it's that testing everything against one real project changes your own relationship to the tools. Instead of a vague sense that "Figma has a lot of AI now" and a fear of missing out on all of it, you come away knowing precisely which two or three features earn a place in your daily work, which handful are worth reaching for on specific jobs, and which you can ignore without guilt. That clarity is genuinely valuable in a space engineered to make you feel you should be using everything, because it lets you invest your limited learning time where it actually pays off and stop feeling behind on features that wouldn't have helped you anyway. The scorecard, in other words, is as much a tool for your own focus as it is content for your readers, and running it once pays back every time you resist the pull to adopt a feature just because it exists.
Keep the scorecard current
One caveat that's also an opportunity: Figma's AI features change fast, so a scorecard is a snapshot, not a permanent verdict. A feature that scores poorly today on reliability may improve, and new features will arrive that belong on the card. Rather than treating that as a weakness, treat the scorecard as a living document you revisit, because a periodically updated "here's what's actually worth using right now" is more valuable than a one-time review that goes stale, and it gives readers a reason to return. Date your scores, note when you last tested, and update the card as features evolve. In a space moving this quickly, the honest, current answer to "which of these should I use" is a genuinely scarce thing, and maintaining it is worth more than producing a single definitive-sounding piece that's out of date in three months.
Start Monday
If you want to build this scorecard, pick your one real project this week and start with the two features you'd use most, likely the agent and Figma Motion, running them through the project and scoring each on speed, quality, and reliability with evidence. Getting real scores on even two features begins the piece and immediately sharpens your own sense of what's worth using. Then work through the rest one at a time. The finished scorecard, built on one real project and an honest rubric, is exactly the resource this fast-moving, demo-flooded space lacks, and it's the kind of tested, evidence-based content that ranks and earns trust.
Figma's AI features don't deserve a single verdict, and the only way to know which earn a place in your work is to run real work through all of them and score the results honestly. Test them on one project, mark each on speed, quality, and reliability, and you'll trade the confusion of a dozen demos for a clear map of what to use, what to use sometimes, and what to skip.