Case studies

Agentic writing support with skills

I led the design of a framework that enabled Quillbot to rapidly develop and scale specialised AI writing workflows, taking the idea from proof of concept to a reusable interaction and component system.

Visual of mid-design states for various agent skill modules

Writing with confidence

Students, academics, and writers more broadly want to be able to submit or publish their work with confidence, and are increasingly turning to unrefined applications of AI to give them reassurance that their writing is polished, professional, and authentic. There was (and still is) ample opportunity here for us to deliver both targeted and end-to-end workflow solutions, and the rise of AI agents presented a way to achieve that.

I lead the design of Quillbot’s writing studio suite of tools, and in late 2025 our team was tasked with envisioning what the future of AI-powered writing support looked like beyond our existing capabilities.

Around the same time, the agent skills standard solidified, and with an AI Chat already sitting inside the writing studio, we began to explore ways in which new agentic capabilities could be authored in natural language, speeding up the release of workflow-based writing features.

Skills to support writing

We believed that skills could offer a way to expand our writing support to benefit not just our core base of students, academics, and educators, but journalists, editors, job seekers, and content creators too.

Throughout the first half of 2026, I designed the agent skills framework: from shaping our understanding of the opportunity space, through iterative prototype testing and development, mapping out conversation flows, designing a new component library for AI agents to dynamically populate, and vetting the text-based skill files that guide AI agents through conversational tasks such claim verification, document grading, literature reviews, and more.

Early agent skills prototype for user research, appropriately janky.

The first skill, claim verification, reached around 7,000 daily users in a matter of weeks, and it taught us some expensive lessons about designing AI products that are useful, reliable, and commercially viable.

The bet

The bet was built on two foundational hypotheses: that we could create a framework that would enable us to rapidly prototype, test, and scale new capabilities ‘faster than we could build them’; and that continuous discovery would enable us to fine-tune skills and establish defensible user value that would be difficult for competitors to copy.

We started with a proof of concept prototype. It was a little clunky, but my PM Kay Brouwers (the real brains behind this whole enterprise) presented to leadership at a company offsite, and within two weeks we’d mobilised a cross-functional team, partnering with AI alignment, to chase the bet.

Kicking off

AI helped us leap into the solution space—but now it was time to pull back. Nothing like a good alignment workshop to pin down our core hypotheses and unknowns with varied stakeholders in the room.

Snapshot from assumptions mapping session
Slow artefacts help to remind us what we know we don't yet know...

From there it was into the cross-functional agile work: I took on more responsibility building out the initial prototypes, running rapid iterations of unmoderated tests and feeding our findings back to calibrate our base assumptions and nudge our prototypes forward. UXR ran deeper parallel interviews and surveys (shoutout to UXR lead Shweta Apte), and I continued to share out learnings as we went. These status updates fuelled the project (and kept leadership in the loop as we encountered unforeseen snags).

Snapshot of highlights from multiple rounds of product discovery
I compiled and shared out highlight reels and reports following each round of product discovery.

This is the work machine intelligence cannot dramatically speed up. It’s also one of my favourite parts of the job. I design for humans, and the more time I spend putting my ideas in front of humans, the more material I have to help me iron out contradictions and deliver more coherent and satisfying solutions.

I design for humans, and the more time I spend putting my ideas in front of humans, the more material I have to help me iron out contradictions and deliver more coherent and satisfying solutions.

Because I was validating assumptions continuously rather than making one big bet and hoping, there was little disagreement on the direction. The real challenges came later…

Tools for thought, not for offloading thought

In earlier versions of our prototypes, our agents proposed rewrite suggestions that could be accepted with a single click. I observed test participants checking off recommendations without even reading what had been suggested, and I didn’t like what I was seeing.

Given our core audience is students, already drawn to the shortcuts that LLMs provide (yes, that’s you, ChatGPT and Claude), I was uncomfortable releasing a capability that would help someone skip learning entirely, and I insisted we design against that.

What is the point of providing writing feedback if it is not properly interpreted?

We rewrote skill instructions so that the framework's job was to nudge people toward the primary source and reasoning, not just hand them a paraphrase to swap in. And I explored a range of options that prioritised explaining before paraphrasing.

Explorations into suggestion card designs
A small sample of explorations into how to lay out and render suggestion card content.

Seems simple enough? Not when part of our conversion model is built around rewrite entitlements: the more users accept rewrite suggestions, the faster they hit their limit that requires them to upgrade. It would also appear that some people I tested with—despite our best intentions—were simply not interested in the why.

Ultimately, I opted not to deliberately obscure rewrites. There is a fine balance between paternally well-intentioned and patronising. But explanations still took priority.

Spicy robots

We were entering the brave new world of non-deterministic user experiences. I had to accept that I couldn’t control every line of agent dialogue, and that the UI components I designed would need to be flexible enough to accommodate all manner of returned values and structural decisions.

An example of the clarifying questions UI in full width mode
Considering how the genUI components can scale for lower vs higher density surfaces.

Given that these components would render not only within the writing studio but wherever we chose to deploy AI agents across our entire product platform, I aligned with our design system team to codify a new library for AI elements and generative artefacts: chain-of-thought, clarifying questions, forms, upsell banners, fallback states, etc.

Every pattern I designed had to account for the full spectrum of scenarios in which it could be used. It was important to make trade-offs in order to reduce complexity, maintain accessibility, and ensure legal compliance and ethical integrity—even for future skills that may not be implemented for a while.

The artifact component from the new AI style system
Building out scalable component variants for the new AI agent library

The much bigger challenge was getting agents to reliably follow their skill instructions at all.

Quality assurance, meet evals…

AI-assisted development is deceptively fast at producing something that looks like a sound idea, but the gulf between a prototype and a safe, secure, reliable, and cost-efficient product is enormous.

I spent months with the AI alignment team (shoutouts to Ashley Parker, Hannah Schindelwig, Nina Wieneritsch, Annabelle Boyer, and Lucie Steible, among others…) and engineers stress-testing every flow two ways: varying the source text itself (does a fact-checking flow that works on a short LinkedIn post also hold up on a 20,000-word essay?), and actively trying to break the conversational UI.

Our test criteria covered bias and identity-based discrimination, harmful or psychologically damaging outputs, privacy and personal data exposure, and more.

The evolution of the clarifying questions flow
The evolution of the clarifying questions flow, which involved repeated rounds of copy refinement with our legal team to ensure we were open about how results are generated.

Although most of these considerations are built into the model provider’s guardrails, it is nonetheless our responsibility to fully evaluate the nondeterministic conversational systems we’re putting in front of millions of people.

The gulf between a prototype and a safe, secure, reliable, and cost-efficient product is enormous.

Thankfully, our engineer Tobias noticed that someone had cranked the model's temperature setting up to what can best be described as “spicy” while testing something unrelated, and it never got reset. A small and silly mistake in hindsight, but an expensive and valuable lesson in how to keep misbehaving systems in line.

The new frontiers of design

AI was already upending how I approach the design process but this project was the first time I fully embraced an AI-assisted workflow.

‘Slow’ artefacts help cover the bases for better, collaborative decision making: mapping out the conversation trees in low fidelity helped us align with engineering, legal, and AI alignment on scope and priorities.

I do not believe that Figma is dead. Far from it. I still found it an essential surface to happy-path conversation trees, pixel-fit the new AI components, and articulate my decision-making regarding the broader system with collaborators. Visualising flows on a canvas makes it far easier to spot what you’re missing.

Visualising flows on a canvas makes it far easier to spot what you’re missing.

But a much larger share of the UX and interaction-model thinking now happens by iterating directly on a working prototype. I either operate directly in our codebase or spin up sandboxes to test more abstract ideas.

Example of a recent coded prototype to explore a new custom rewrite field.

Slow is smooth; smooth is fast

I wouldn't say any of this makes the process faster. Designers are often seen as pesky bottlenecks, and in my opinion, bringing AI into our practice will do little to alter that reality.

Good product design necessarily requires a certain cadence. It demands that certain considerations be taken into account, and whilst LLMs can certainly produce more, I’ve seen little example of that mapping to increased efficiency.

Snapshot of a branch of our Opportunity Solution Tree
Artefacts like Opportunity Solution Trees help keep us grounded (or rooted, you might say...) to the real outcomes our designed solutions strive for.

The old, slow questions have not gone away: Who are we building this for? Why do they need it? In what ways can we provide value that others overlook? Would they realistically be willing to pay for it? Does it align with our broader company strategy? Is there a innovative interaction model that can act as a moat?

Perhaps the only question whose answer has nominally changed is “Is this technically feasible?”.

Scaling the system

We rolled out the claim verification skill throughout July, and by mid-August it was being used by ~7000 unique users each day. Given the cost of running skills, we decided to gate them behind a paywall, but what we’d failed to consider was that even premium subscribers could run up a tab (the top scorer was one user who managed to run $32 of claim verification scans)!

Without additional usage restrictions (something we’ve historically tried to avoid through caching and optimising LLM calls) our new skills library would fail to cover its own costs of operation. So, we aligned and put new daily/weekly limits in place.

Another learning, another design pattern.

Moving to modules

Over repeated rounds of user testing, I observed the same patterns: no one cared to intervene using the chat input field, and users complained about verbose responses, slow wait times, and too much information being rendered in a cramped space.

Whilst AI Chat forms the technical backbone of these conversational workflows, I never believed it was the ideal interface, which is why we’re launching standalone modules that strip out some of the LLM ‘fluff’ and position skill workflows alongside adjacent capabilities, such as plagiarism checking and AI detection.

The final (static) design of the integrated fact-checker skill
The final (static) design of the fact-checker (formerly claim verification) agent skill integrated as a standalone module in the writing studio.

At the time of writing (September 2026), claim verification has been in beta for two months, and we are about to roll out the second skill: AI grading (literature review, resumé review, improved plagiarism checking, and more are in the pipeline).

The discovery work keeps evolving too, moving from abstracted prototypes toward direct in-app testing with real users, in context, as more of this ships for real.

Designing for acquisition thru retention

Each new skill will be paired with landing pages to improve top-of-funnel SEO/GEO authority, and every improvement I make to the design and performance of various touchpoints maps consistently across every new capability. We can already see the compounding benefits of this in product analytics: customers come to detect plagiarism and stay to verify claims, or vice versa.

It’s in this ‘cross-pollination’ between different modes of writing support that the writing studio excels and becomes increasingly indispensable.