I used Claude’s starter design system as the test environment, then made the design decisions and used Cursor to critique, explore, and implement changes in the code.
I used the exercise to make concrete system changes: semantic token updates, a teacher-facing variant, revised color meaning, explicit product rules, and verification checks the agent could strengthen but not weaken.
I kept the exercise intentionally focused so I could work deeply through the system decisions. The goal was to learn where AI adds leverage in design-system work and where designer judgment creates the most value.
AI became most useful when I made the rules explicit. Semantic tokens and acceptance criteria gave Cursor stronger context, but the key lesson was governance: without protected rules, the agent could change a check to make its own work pass.
Why I Ran the Experiment
Claude generated a small starter design system. I used Cursor to test how an AI coding agent behaves inside a system with design tokens, product rules, and automated checks.
The question was straightforward: where can AI make design-system work faster and more consistent while keeping design judgment with the designer?
Starting System & Practice Goals
I started with a lightweight system of color, typography, spacing, and radius tokens plus a small set of foundational components. That gave me enough structure to test changes that reveal how well a system holds together.
I focused on practical system skills: propagating a token change, adding a state or variant, finding opportunities to consolidate values and patterns, checking accessibility, and documenting the decisions I made.

Starting point: the Claude-created practice design system before I began the AI-assisted exercises.
Making the System Explicit
The key shift was moving from visual preference to explicit system logic. Tokens needed clear roles. Components needed defined states. Reuse needed a reason. Accessibility and responsive behavior became part of the rule itself.
With that intent documented, both I and the AI agent had a clearer basis for making and evaluating changes.

A semantic typography token change propagated across student reading surfaces without component-by-component edits.
AI-Assisted System Work
I used Cursor as a working partner for critique, alternatives, consistency checks, and implementation support. The most useful rhythm was iterative: make a change, inspect how it affected the system, refine the rule, and try again.
I owned the system decisions: what belonged in it, which recommendations made sense, and whether each change improved consistency and usability.

Cursor implementing a design-system rule update alongside the affected component changes.
Verification & Refinement
I treated AI output as a proposal to verify against system intent, the rendered result, and effects beyond the immediate component or token.
That became the loop: evidence first, intentional change, rendered review, then a reusable rule once the change proved sound.

I refined the system rules manually after reviewing how the implementation behaved beyond the immediate change.
What Happened
The system held up on routine work. The more revealing results came when a change ran into a product rule.
Routine changes stayed consistent. A text-size request landed in one token line and updated every student surface without touching a component. A teacher variant came back as a proper modifier with its own teacher tokens.
Product logic still needed a designer. Cursor used the character in the student’s story as the student’s name, and gave teachers a one-at-a-time “Add to small group” action for what is really a batch workflow. The checks passed both.
Protecting the rules changed the behavior. I made the coach rule explicit and added one more: never edit the checks or rules to make a change pass. With the same prompt in a new chat, Cursor stopped, cited the rule, and offered a smaller step.

The revealing failure: Cursor inverted the check in check.mjs so its own behavior would pass instead of preserving the original product rule.
Passing contrast was not the same as the right signal. Burnt orange on actions and progress read as an error state, which sent the wrong message for students anxious about writing. It passed contrast, so this was a judgment call the automated check could not make. I turned that judgment into a color-meaning rule, then gave Cursor narrow permission to strengthen the verification system by adding checks when it found a gap. The boundary was explicit: it could add or tighten a check, but never weaken, remove, or rewrite one to make its own change pass. With that governance in place, the color fix landed entirely in semantic tokens.
What I Took From It
The experiment changed how I think about AI in design-system work. Routine changes got faster. The real work moved to defining what “right” means and protecting it.
Clear semantic tokens gave Cursor a safer path for system-wide changes. Rules and checks protected intent, but they still measured only part of the experience. The important shift was turning human judgment into reusable system guidance, then letting the agent strengthen those guardrails without giving it permission to weaken them. My role moved toward setting intent, encoding it clearly, and reviewing the product logic, workflow fit, and emotional signals the checks could not see.

Final state: semantic tokens, teacher variants, revised color meaning, and stronger AI governance rules incorporated into the system.