June 3, 2026
Generate, Evaluate, Refine: Design Judgment in AI-Assisted Product Design
There's no shortage of AI workflow advice for designers these days, and most of it stays abstract. Just open LinkedIn.
What follows emerged from a real project. Three phases that took us from an empty canvas to clickable prototypes ready for concept testing in a single afternoon. More interesting than the speed, however, was the role design judgment played in each phase.
The workflow: generate, evaluate, refine. Here's how it actually worked.
I spent several years designing software that helps teachers run their classrooms. As the product evolved, the navigation started getting in the way. It couldn't accommodate new features gracefully, and the experience had grown more complex than it needed to be. Design, engineering, and product all knew we needed a navigation that could support the product's next stage of growth while feeling simpler to use.
We also knew we wanted a place to surface timely information throughout a teacher's day, but it was unclear where that experience belonged or what it should contain. The navigation had to come first because understanding how teachers move through the product was a prerequisite for knowing what a time-based experience should surface.
Our timeline was tight. We needed clickable prototypes ready for concept testing before the next design review. What we didn't know going in was that work we expected to take days would compress into a single afternoon, and that AI would play a central role in making that possible.
The research behind the sprint
To kick the project off, we ran a moderated open card sorting exercise with academic staff and product team members. The goal was to understand how teachers think about their daily work so we could design a navigation that better reflected their mental model.
Kicking the project off with a 15-minute open card sort exercise.
The results were remarkably consistent. The card groupings and category labels varied very little across participants, giving us a clearer picture of how teachers organize their day: what they reach for first, what they return to repeatedly, and what they rarely need at all. From those patterns, a new navigation architecture emerged, one organized around how teachers spend their time rather than how the software happened to be structured.
The card sorting phase matters for one reason. We already knew this product well. Years of working on its features meant we understood how teachers used it and where it was failing them before the card sort began. What the exercise produced wasn't new product knowledge; it was a clearer understanding of how teachers think about their work.
With the navigation architecture in place, we could finally turn our attention to the time-based experience. The question wasn't simply what information teachers needed, but what information they needed at that moment in their day.
Surfacing the right information at the right time
The first design review focused on navigation. We proposed a sectioned, icon-anchored structure that reduced the number of visible options at any one time. The time-based experience, however, remained an open question. We knew it belonged in the product, but not what information it should surface or how it should work.
To answer that question, we explored two concepts. The first was a dashboard containing information relevant to every teacher, regardless of grade level or subject. The second was a shortcuts feature that let teachers pin the screens they used most often. This wasn't simply an alternative concept; it was a deliberate risk-reduction strategy. Rather than trying to predict every widget teachers might need before the school year began, the shortcuts feature ensured teachers could always surface what mattered most to them. If we got a widget wrong, shortcuts would fill the gap.
By the end of the review, we had alignment on both concepts and an initial set of dashboard widgets that every teacher would need. The list reflected the moments that shape a typical school day: current class, attendance, unread inbox messages, grading progress, parent communication, upcoming events, academic reporting, and seasonal workflows like re-enrollment and academic progress testing.
That list left the room with us on a Thursday morning. By the next design review, it needed to exist as a clickable prototype ready for concept testing.
IA diagram with widget ideas that emerged from first design review.
A dashboard isn't a screen. It's a snapshot in time.
Designing a dashboard isn't about arranging cards on a page. Every widget has to tell the truth about a particular moment in a teacher's day.
Attendance means something different before 9:00 AM than it does after it's been submitted. The current class widget needs different information at the beginning of class than it does five minutes before the bell. Grading changes throughout the day as work is completed, while parent communication shifts as new messages arrive.
Clickable prototype with time aware widgets for testing.
Getting that logic right across fifteen widgets, each with its own states, rhythms, and relationship to the clock and calendar, is what concept testing ultimately validates. That's the work that occupied the afternoon. Here's how it happened.
Generate, evaluate, refine
With fifteen widgets to design and an afternoon to do it, the process couldn't afford to be slow. What emerged wasn't simply a faster way of working with AI. It was a workflow where each phase had a distinct purpose, and where the role of design judgment changed throughout.
Generate expanded the solution space.
Evaluate applied product knowledge and design judgment to identify what belonged.
Refine transformed promising concepts into product experiences through an ongoing dialogue between designer and AI.
Each phase has a distinct role. Here's how they worked in practice.
Generate: Expanding the solution space
The first move was to open a conversation with Claude and start exploring each widget without constraints. I deliberately avoided loading product context, referencing existing UI patterns, or connecting Figma MCP. The goal wasn't accuracy. It was volume and variety: as many interpretations of each widget's content and structure as possible before making any design decisions.
For each widget, I gave Claude a focused prompt that established the problem without prescribing the solution. I started with the current class widget:
Provide options for a current class widget that communicates time remaining, time elapsed, when the class ends, and uses color to provide quick visual reference. What are the different ways you can lay out the content for this widget?
Claude generated class widgets.
The deliberate decision not to overcontextualize at the start mattered. When AI knows too much about your constraints upfront, it begins to self-edit, producing safer, narrower output within the lines you've drawn. Leaving those constraints out expanded the solution space. Most of the concepts would eventually be discarded, but they surfaced in minutes. A paper-and-pencil sketch session that might have taken an hour was compressed into five minutes.
The generate phase is intentionally high-volume, low-stakes thinking. Like a traditional design charrette, the goal isn't to find the answer. It's to explore the breadth of possible answers before deciding which ones deserve further consideration.
The constraint is no longer AI's ability to generate ideas. It's the designer's ability to evaluate what comes back. That's the next move.
Evaluate: Design judgment creates value
The generate phase returns possibilities. The evaluate phase is where the designer determines which of those possibilities belong in the product. This is where design judgment creates the most value, not because AI has failed, but because evaluation depends on knowledge that is difficult to transfer into a prompt. It requires an understanding of the product's interaction philosophy, the experience principles that have evolved over time, and the deeper logic that holds the product together. Those things don't exist in isolation. They're accumulated through years of designing, shipping, and refining the experience.
The dashboard project surfaced two different kinds of misalignment. The first is the easiest to recognize: AI adds something that doesn't belong. For the attendance widget, Claude suggested introducing a streak counter to encourage teachers to complete attendance by 9:00 AM each morning. It was a thoughtful idea. Habit tracking works, and in another product it might have been the right solution.
Claude introducing streaks into the product.
The problem wasn't the idea. It was the product. Nothing else in the experience used gamification, and introducing it in a single widget would create inconsistency without a broader product decision to support it. The streak wasn't removed because it was a bad idea. It was removed because it didn't belong.
The second type of misalignment is more subtle. Rather than adding something new, AI introduces an interaction pattern that appears reasonable but quietly contradicts an established design principle. One of the dashboard widgets summarized communication between teachers and parents. Claude responded by designing a complete posting interface directly inside the widget. The interface looked polished, and on its own it was a perfectly reasonable solution.
Claude expanding on communication UI into a fully functional widget.
The problem was that it violated one of the dashboard's core design principles. The dashboard had been intentionally designed as a surface for awareness, not action. Its purpose was to summarize information and direct teachers to the right place when action was required, not become another workspace within the product. That's what makes this kind of misalignment more difficult to recognize. The interface isn't wrong. The interaction model is.
Both examples point to the same conclusion. Wrong additions often feel out of place. Wrong interaction patterns can feel completely natural until they're measured against the principles that define the product. Both require design judgment. Once the strongest concepts survived evaluation, it was time to refine them.
Refine: Designer and AI in dialogue
The third move was to take what survived the evaluate phase, design the interface, and bring screenshots back to Claude with specific questions. It's the move that's easiest to skip because the concepts AI generates often feel complete. In practice, we found it to be the most valuable. The current class widget illustrates why.
During the generate phase, Claude proposed three approaches for communicating time in a class: a horizontal progress bar, segmented blocks, and an arc with inline metadata. The evaluate phase selected the horizontal progress bar as the strongest foundation. The other concepts were set aside, but selecting a layout wasn't the end of the design process. A progress bar only tells a teacher how much time is left. The refine phase is where the designer begins adding what AI doesn't know.
For our dashboard concept, we decided the current class would be literacy. That decision immediately introduced context the AI couldn't infer: the book being read, the pages assigned for the day, the current lesson, and a direct link to the teacher's notes. None of those details came from Claude. They came from understanding how a teacher's day unfolds, what they need at 12:02 PM when one class ends and another is beginning, and how little time they have to search across the product before students arrive.
Refining the class widget with real class information.
That richer design went back to Claude as a screenshot with a simple question: Here is a design of the current class module with additional details. How can I refine this experience? Give ideas that have time context. The conversation immediately shifted from interface design to experience design. Instead of generating another layout, Claude began reflecting on the experience itself: whether the widget surfaced the right information at the right moment, whether anything was missing, and how it might better support teachers as they transitioned between classes.
The conversation shifted from generating interfaces to reasoning about product behavior.
One suggestion stood out. Five minutes before the current class ends, a teacher is already thinking about what comes next. Should the widget preview the upcoming class before the transition happens? Could teachers customize the information they see during those few minutes between classes? Those weren't UI questions. They were product questions, and they emerged because the conversation was grounded in a specific design informed by product knowledge rather than an abstract prompt.
That's what distinguishes the refine phase from the generate phase. Generate explores possibilities. Refine builds on product knowledge, then uses AI as a thought partner to challenge assumptions, identify opportunities, and strengthen the experience. The conversation is no longer about what the interface should look like. It's about what the product should do for the person using it.
What made it work
One thing is worth noting before drawing any conclusions. The sequence, generate, evaluate, refine, isn't always linear. The examples in this article move cleanly from one phase to the next, but in practice the workflow is far more fluid. You might be deep in the refine phase, bring a screenshot back to Claude, and discover that the right response is to generate an entirely new set of concepts with the additional context. Those concepts are evaluated, refined, and sometimes sent back again. The three phases are better understood as modes of thinking than a fixed sequence of steps, and knowing which mode the work requires at any given moment is itself part of the design process.
What made this workflow successful wasn't AI. It was the foundation that existed before AI became part of the process. Years of working on the product meant we understood its interaction philosophy, the design principles that had evolved over time, and the countless small decisions that gave the experience consistency. That knowledge is what made the evaluate and refine phases possible. Without it, the concepts generated by AI all appear equally plausible because there's nothing meaningful to evaluate them against.
AI amplifies what the designer brings to it. Feed it product knowledge, design principles, and a deep understanding of the people you're designing for, and it becomes an extraordinary exploration partner. Without that foundation, the output may look polished, but it fits nowhere in particular. It's simply interface.
The honest version of the project is straightforward: fifteen widgets, a single afternoon, and a clickable prototype ready for concept testing. Those outcomes are real. But they weren't the result of AI working independently. They were the result of knowing which ideas to pursue, which to discard, and when to move beyond generating interfaces to reasoning about the experience itself.
The pace of design has changed.
The role of design judgment hasn't.