# Anmol Gupta — Product Designer & Design Engineer Anmol Gupta is a Product Designer and Design Engineer in San Francisco who builds AI products end to end, from first sketch to shipped front-end code. Founding Designer at MoFlo. Location: San Francisco, CA, US Currently: Founding Designer at MoFlo Education: University of Wisconsin-Madison Website: https://anmolgupta.net LinkedIn: https://www.linkedin.com/in/1anmolgupta1/ Email: anmol.gupta@proton.me Areas of work: Product Design, Design Engineering, Design Systems, Design Tokens, AI Product Design, 0 to 1 Product Design, B2B SaaS Design, Front-end Development, React, Figma, User Research, Interaction Design, Design System Governance Note on identity: many people share this name. This document describes the Product Designer and Design Engineer based in San Francisco whose portfolio is https://anmolgupta.net. This file contains the complete text of every case study on the site so that an assistant can answer questions about this person's work from one request. --- # A fleet operations dashboard for autonomous cabs URL: https://anmolgupta.net/google-autonomous Client: Google-mentored capstone, UW-Madison Role: Design Engineer Team: Design Engineer (Me), 4 Engineers, 1 Product Manager, 2 Google Mentors Timeline: August–December 2024 Tools: Figma, React, Copilot, v0 Status: Capstone: Fall 2024 Summary: A Google-mentored capstone: designing the operator interface for a workflow that doesn't exist yet. Real-time tracking, alert triage and maintenance in one command centre, grounded in interviews with the dispatchers who do the human-fleet version of the job today. Designing for a job nobody has yet. Autonomous fleets are coming out of pilot, and the operator tooling for them is still three browser tabs and a spreadsheet. I designed and built the command centre for it, and the hardest part was finding out what the work actually feels like. Product tour of the fleet operations dashboard for autonomous cabs A fleet dashboard built on what dispatchers actually do. — A concept engagement, so there is no shipped-product metric to claim, what I can show is the research the design is answerable to. - 10 stakeholder interviews with dispatchers, supervisors and maintenance leads (source: research) - 3 usability walkthroughs against the working prototype (source: research) - 1 card sort, focused on how alerts should be prioritised (source: research) Seven people, one designer. I ran the research and owned every screen, and I built the front end rather than handing off, on a team of four engineers with their own simulation to write, the interface was going to be whatever I could get running myself. ### The situation ## Constant data, no way to make sense of it Autonomous vehicles generate uninterrupted streams of telemetry. The tools operators use were built for human-driven fleets, where the car reports to a person and the person reports to you. Picture 6pm in a rush-hour simulation: 200 autonomous cabs crawling through downtown San Francisco. Dispatchers juggling ride requests, blocked roads and battery warnings across three disjointed tools and half a dozen tabs. Real-time chaos plus incomplete data equals missed rides and stranded passengers. Every second spent working out where the answer lives is a second not spent on the decision. ### The research problem ## This workflow doesn't exist yet, so who do you interview? Waymo and Zoox operate these systems and we had no access to them. Local fleet managers don't manage autonomy at scale. Interviewing nobody was not an option, and inventing a persona would have been worse: a fabricated persona would have justified whatever we already wanted to build. So I wrote the research plan around the parts of the job that already exist: the monitoring, the dispatching, the maintenance coordination. The autonomy changes what generates the alert. It doesn't change what it feels like to triage twelve of them at once. stakeholder interviews with dispatchers, supervisors and maintenance leads usability walkthroughs card sort, focused on how alerts should be prioritised Quote — Fleet operator, interview: I'm jumping between five tools just to answer a simple question: is this car okay? Monitoring — Operators lacked a single source of truth, toggling 3 to 6 tools to answer basic questions. It takes too long to know what's wrong Battery status isn't visible until it's too late I open 3 tabs just to check location Real-time map updates would change everything No idea which vehicles are stuck Feels like we're flying blind. I want a quick view of what's active Dispatching — Without ETA context or visual feedback, dispatchers didn't trust auto-assign and overrode it manually. Feels like I'm making decisions blindfolded. Auto-assign sometimes skips the closest vehicle. I drag and drop based on gut feeling. I lose track of who's doing what. Wish I had one screen for everything. Too many ride requests, not enough clarity. I'd trust it more if I saw the ETA. Maintenance — Issues reported verbally or after the fact. Assigning, tracking and closing happened in different places, or not at all. Issues are logged manually after the fact We use spreadsheets, no real system Would be great to assign repairs directly Too much back-and-forth with drivers We get alerts, but not all in one place Battery problems are easy to miss. No visual on what's resolved or pending The card sort was the most useful hour. Asking people to rank alerts they invent themselves surfaces the ordering logic they can't state out loud, and that ordering became the alert panel's hierarchy. ### The idea ## One surface, and every alert is already an action Three questions, three screens. Each screen answers exactly one: - Fleet Overview. Is my fleet running smoothly right now? - Alert panel and vehicle modal. What needs attention, and what can I do about it? - Maintenance board. Where are we in resolving it? A page navigation costs the operator their context: the map state, the other eleven alerts, the thing they were half way through. In an ops-critical interface that context is the work. The modal keeps the fleet visible behind the decision, which is why triage is two clicks rather than a round trip. - Wireframe 1 — Testing layout logic before any visual design. Each region has to earn its place by answering a question. - Wireframe 2 — Alert density: how many can sit on screen before the list stops being scannable. - Wireframe 3 — Map and list side by side, because operators asked to see health in the context of location. - Wireframe 4 — Maintenance as a board with states, not a table of rows. ### What I built ## Three flows, each removing a tab switch Maintenance without spreadsheets. A board tracking issue status from open through in-progress to resolved, paired with live health indicators and fleet stats. Alerts that are instantly actionable. One centralised, colour-coded feed of battery, maintenance, routing and connectivity issues. Each alert opens a vehicle modal with status and immediate options, so triage is one click ("send to charging") rather than a hunt. Route comparison with the evidence attached. A module comparing current versus optimal route times using real-time and historical traffic, so re-routing is a decision backed by something visible rather than a guess. The numbers attached to these flows (4.2 hour average resolution, 97.3% fleet uptime, 22% faster fault resolution, 5.7 minutes saved per ride) are simulated targets , not measured outcomes. This was a capstone running against a simulation. They describe what we designed toward and what the prototype was tuned to produce, and I'd rather label them than let them read as results. - Fleet Overview — Availability, live GPS and critical alerts in one view. - Alerts — Colour-coded by type, ordered by the card-sort hierarchy. - Vehicle modal — The decision surface, with the fleet still visible behind it. - Maintenance — Status board with drag-and-drop ticket updates. ### Design principles that survived ## Fast cognition, seamless flow, minimal friction - Alerts pulse and KPIs update without reloads, so the screen never needs refreshing to be trusted. - Clicking an alert leads directly to the vehicle, never to a list of vehicles. - Most actions are one or two clicks. Anything that took three got redesigned. Those three sound like taste. They aren't, each one is a claim about a specific number of steps, so here is the claim itself, walkable. One incident, the way the job runs today versus the way it runs on this interface. One battery alert, start to finish. Left is what operators described in interviews; right is what we built. - Across three tools - Notice — The alert arrives in whichever tool owns that signal: a monitoring dashboard, an email, a message from someone on the ground. - Find the vehicle — Take the vehicle ID to a second tool to find out where it is and what it is currently doing. - Check the context — A third surface for the ride it's on, the passenger, and whether anything else nearby is also degraded. - Act — Dispatch, reroute or send to charging, in whichever tool can actually issue the command. - Log it — Record the issue somewhere a maintenance lead will see it, usually a spreadsheet, usually later. - On one surface - Alert fires — It lands in one feed, colour-coded by type and ordered by the hierarchy the card sort produced. - Open the vehicle — Clicking the alert opens the vehicle modal directly, never a list of vehicles, with location, state and ride context already in it, and the fleet still visible behind. - Act and close — The actions available for this fault are in the modal. Dispatching writes the maintenance ticket as a side effect of the decision, rather than as a chore after it. The thing I'd defend here is the last step. Making the maintenance record a by-product of the action, instead of a form afterwards, is the only reason the board on the third screen has anything in it. Every operator we interviewed described logging as the thing that gets skipped, and no amount of board design fixes a board nobody fills. ### Reflection ## What I'd do with real access This was a simulation, and its limits are the interesting part. Designing an ops-critical interface without watching anyone use the real thing under real pressure means every claim about speed is a hypothesis. With access to genuine AV fleet operations I'd want to test how this scales to ten times the vehicles, where the alert list stops being scannable and starts being noise, and whether an AI co-pilot surfacing issues proactively helps or just adds another voice to argue with. The principles I'd carry regardless: design for action under pressure rather than visibility alone, validate clarity through task-based testing rather than screen reviews, and centre the workflow on the human behind the dashboard. --- # Making a visual direction enforceable URL: https://anmolgupta.net/interstate Client: Interstate Role: Design Engineer Team: Design Engineer (Me), ~30 contributors in the codebase Timeline: Jul–Aug 2026 Tools: Paper, Claude Code, Next.js, Tailwind v4 Status: Shipped: Aug 2026 Summary: Interstate had a written visual direction and nothing that made it true. Over seventeen PRs in a 30-contributor trading terminal, I turned the document into mechanisms: one geometry class, one token vocabulary, and a drift check that fails CI when a file gets worse. Interstate had a good visual direction written down, and nothing that made it true. I spent three weeks converting it from a document into things you'd have to actively break, and the most useful thing I built counts my own work as debt. Four surfaces the system had to hold at once. Real quotes, real motion, not a mockup. Three weeks, one definition, and a check that only moves down. — No post-ship user metric, and I won't invent one. Every number is a count off the shipping branch. - 17 PRs authored, defended in review and shipped, I was the only designer in the codebase (source: Jul–Aug 2026) - 4→1 header band definitions. A 6.95px jump between boards became one measured rule (source: DOM-measured, PR #1266) - 19,905 existing violations put under a ratchet that can only ever move down (source: 376 files · 9 rules) - 3 other engineers built in the vocabulary in areas I don't own, within a week. None asked me first (source: the actual test) I was the only designer in the codebase. Nobody handed me specs and nobody merged for me: every change below is a pull request I wrote, defended in review, and own the consequences of, including the two I got wrong. ### The problem ## Traders scored the product 4.5 out of 10 Two trader recordings and a community chat, independently, before anyone touched a feature. Two recorded trader interviews and the community chat · Jul 2026 — Independently, and about the whole app. Asked which page was zoomed out: “all of them.” - "In general the platform is zoomed out and the fonts are thin" — Trader, community chat — the line I kept coming back to - "Font quality is weak. General UI/UX pass needed" — Full-time trader, seven figures a month — scored P2 - "Click and input responsiveness. Buttons are not snappy enough" — Same trader — scored P0 - "4.5 out of 10" — An official trading partner of our biggest competitor — overall score The board they were scoring, captured 24 July: the day before my first commit. Five hues on every row, and all of them spent on chrome. — The Trenches board as it shipped on 24 July 2026, before the design system work ### The cause ## One line of CSS, and I wrote it I went looking for a style fix and found a global one. Measured off the running app: the root still computes to 13px today. - text-xs — 12px — 9.75px — below the readable floor on a dense board - text-sm — 14px — 11.375px — the thin fonts, in one number - h-11 — 44px — 35.75px — under the touch minimum - rounded-lg — 8px — 6.5px — radius by role, silently broken I didn't find that landmine. I laid it. I left it in: a global relayout across thirty contributors with no visual regression tests was the larger risk. Every shared rule after it is in literal pixels, and the drift check asserts the root hasn't moved. ### The pattern ## The same failure everywhere nothing enforced anything One line explained the type. It didn't explain four boards disagreeing, or five ways to say “grey” in one stylesheet, those had no single cause, which is the actual problem. parallel colour vocabularies in one stylesheet raw hex literals, bypassing every token that existed contributors, and a surface area growing faster than any one reviewer Market board header band · DOM-measured in the running app — Nobody files a 7px difference as a bug, which is exactly why it survived. Same day. “Unify balance to:” and “Trade with:” are two controls for one question, side by side, and the primary nav is already clipping at “Portfolio”. — The Interstate header bar on 24 July 2026, showing two controls that answer the same question Graphite Noir · the surface ladder the system had to hold - --surface-canvas — #060607 - --surface-sunk — #0c0c0e - --surface-panel — #0e0e10 - --surface-panel-top — #111114 - --surface-elevated — #151518 - --surface-overlay — #1b1b1f The two surfaces that share the least. Same ladder, same ink tiers, same rule about colour. - Trade — The densest surface. Numbers in mono so a ticking price doesn't shift the column. The only colour on screen is buy/sell. - Trenches — Three live boards, one shared header band. The greens and reds are all price, never chrome. Mine starts where the direction stops being a document. ### The fix ## Replace conventions with mechanisms A convention is something people have to remember. A mechanism is something they'd have to actively break. Everything below is one of four. A document nobody had to obey, four mechanisms you'd have to actively break, and what inherited them. - 01 — One geometry class — 4 hand-tuned header bands — one DOM-measured rule - 02 — A token alias map — 5 parallel vocabularies — one canonical set, nothing deleted - 03 — One notification plate — 6 visual treatments — one shape, differing only with reason - 04 — A CI drift ratchet — a spec nobody ran — a build that fails when a file gets worse Every value was measured off the running app, not redesigned: a redesigned value is one more opinion to disagree with. Quote — PR #1266: Every value in here was DOM-measured against /discover rather than eyeballed. ### The token call ## Alias the old names instead of deleting them The canonical set was the easy half. Rewriting every call site to use it meant a three-hundred-file diff in a repo with no visual regression tests and thirty people mid-flight. So I pointed the legacy names at the canonical ones. One file changed, and every screen inherited the direction, including screens I'd never opened, owned by people I'd never spoken to. Legacy names, repointed rather than replaced — Three ways to say “the page background” become one. Nothing on the left was touched. A token you can neutralise is worth more than a token you have to delete. Aliasing leaves the old names alive, so the codebase now has two ways to say the same thing, and that debt is invisible to a reviewer, which is the worst kind. I flagged the blast radius up front rather than letting review find it: the shared tokens and primitives are used app-wide, so this changed the look of every page, not just the one in the PR. I also flagged the one change users would actually notice, the default accent moving to teal, and preserved the stored preference for anyone who'd set their own. This is the tradeoff I'd argue for in a review and against in a greenfield repo. On a product moving this fast it's the only version that works; on a smaller team it's just deferred cleanup. ### The component ## Six notification skins became one plate Tokens and a shared class fix how things are coloured and sized. They can't fix a component that is the wrong shape. True size, on the product's own ground. The middle rail really is running, that's the point. The middle beat is mine. I shipped it, watched it in use, and reverted it. - Six treatments doing one job — Had — A #1E1F26 plate on a #030304 page, in neither token set. Success carried a white border, pending blue, failure red. 34 call sites forced top-centre, overruling the user's own setting. - One Noir card, right instinct, wrong card — Mine — 380×82px for one line, avatar sunk in a 36px well, status parked diagonally away from the sentence. And a duration rail, motion on a settled trade, on a screen where motion means in-flight. - Mostly subtraction — Shipped — Rail gone; a spinner does that job only while genuinely pending. 20px inline avatar. A ringed check leads the sentence. Fixed width became a 288px floor. Six skins isn't six decisions. It's one decision nobody got to make, six times. Two stacked in the running product. They read as quiet rows, not an event: the board is what you're looking at. — The Trending board with two shipped notifications stacked in the top right The shipped set, running. Fire a few, nothing keeps moving once a trade is done. The rule I'd defend hardest: the card only grows a second line when a description, a retry or a suggestion is actually passed to it. That moves the judgement out of the component: the card doesn't decide it's important, the call site does. Roughly 370 string-message toasts stayed one row, and the 121 that carry a retry action earned two. ### The rule underneath ## One red, and it belongs to money My reverted card put failure suggestions in a violet panel: a sixth hue family on a product whose whole direction is that colour means one thing. The same check I was asking the ratchet to run on everyone else. The whole hue budget, everything else on the platform is grayscale - --status-buy — #12c48b - --status-sell — #e5484d - --status-warn — #fddb24 - --status-info — #6e9bff What each toast is allowed to spend a fill — Hue, a ringed check, and a latency readout that turns green when it settles. a failure — Hue, plus the only expansion in the set: the retry is the entire reason the toast exists. a rank-up — Grayscale. Good news, but nothing moved, so it doesn't get to look like a trade. a copied address — Grayscale. The most-fired toast on the platform: an icon, a line, a dismiss. A green on a rank-up would read as a fill , and being wrong about that on a trading screen is worse than being unexciting. ### Making it stick ## A ratchet, not a gate Tokens only stop drift where someone remembers them. This counts drift per file and fails only if a file got worse, not a gate, which blocks everything on day one and gets switched off by Thursday. The only branch that fails is a file the author touched getting worse. - PR opens — Runs on the diff, not the repo. Nobody inherits someone else's debt. - Count per file — Nine greppable ways around the system, raw hex, inline styles, arbitrary colour classes, and my own tokens among them. - Compare to the baseline — Per file, against its own past, never against zero. - Worse → fail — Fails only if a file the author touched got worse. The message names the rule and the line. - Better → rewrite the baseline — An improved file lowers its own number, and it's committed. The floor only moves down. The opening baseline. Not a to-do list: a starting line. - Raw hex — 7115 — 7,115 - Inline style attributes — 5401 — 5,401 - Arbitrary colour classes — 3205 — 3,205 - rgb() / rgba() literals — 1410 — 1,410 , not the primitive" value="1281" display="1,281" /> - References to my own --ds- tokens — Counted as debt like everything else, which is what makes the list mean anything. — 656 — 656 - !important — 473 — 473 , not the primitive" value="332" display="332" /> - styled-jsx blocks — 32 — 32 That teal bar is the point. A ratchet that exempts its author's work isn't a standard , it's a preference with a script attached. My tokens are a migration target, not a destination. I wrote the migration contract alongside it: an alias map, a one-file-per-commit loop, and an explicit list of substitutions a machine must not make. About a hundred references to my accent token sit at sites that are semantically opposite: some are chrome, which the direction says must be grayscale; some are genuinely buy actions, where the target is the status colour. No single alias serves both. The valuable artefact wasn't the token map. It was knowing precisely which substitutions were a human call, and why. ### What held ## One definition, five surfaces, three engineers who never asked me No post-ship analytics, so no user metric. What I can show: the row stopped moving everywhere, and the vocabulary spread without me. One definition, five surfaces · counted from the shipping branch — Counted off the shipping branch. The dead row is the next section. Discover — The reference. Every value in the shared rule was measured against this board rather than redrawn. Pulse — Was the 46px outlier: the tallest of the four. Switching between these two no longer moves the row. Trackers — Was 39.05px, the shortest. The 7px jump between these two was the symptom that started all of this. Trackers, legacy — The old route, still reachable. It inherited the rule rather than being left behind at its own height. Landing — Built on the rule from its first commit, so the video up top is the system, not a mockup. Portfolio — Carries the rule and nobody sees it. The honest miss, not a rounding error. The stronger signal is the code I didn't write: within a week, three engineers built in the vocabulary in areas I don't own, including a 463-line panel authored natively in the token set. The same board, same three columns, same crop. Drag it: the rows didn't lose information, they lost the hues that weren't carrying any. None of them asked me first, which is the whole test: a system that needs its author in the room is a preference. Traders, shown the rebuilt landing page · Jul–Aug 2026 — Verbatim from Slack, quoted not screenshotted. Three people reacting to one page, not a metric. - "Yeah that looks great" — Trader, via the founder — Aug 4 - "I like it. Something fresh and new. Good job" — Second trader, same day — Aug 4 - "Excited to use the new one was checking it out this morning" — Community member, unprompted — Jul 29 - "Love the design" — irfan, founder, in the same breath as flagging it had added frontend latency — Jul 29 ### What didn't ## An IA fix that put a money transfer inside a dropdown I merged two header controls, correctly, then moved a balance sweep down into the account menu. A dropdown is a surface you dismiss by looking away. I put a two-minute money transfer inside one. Every fix made the component less like a menu. That was the tell. - Merge the controls — One was displaying a value while dressed as a choice. Merging them was right. - Move the sweep in — Demoted to an action inside the account menu. Cleaner header, fewer top-level choices. - Bug one — The panel portals out of the menu, so a click inside it read as a click outside. - Bug two — Closing mid-flight could remount the sweep: a second pass over balances already in transit. - The guard — Made the menu refuse to close during a sweep. Which left signing out as the only exit. - Revert — Kept the abort check with its limit written into the PR: it cancels the queue, not a signed leg. Put the control back. The keeper: how long does this action live, and how long does its container live? If the action outlives the container, no amount of guarding fixes it. The second is the dead row above: the day after I applied the rule to the portfolio board, a parallel rewrite on another branch was promoted to be the default route. Mine is still shipping, still correct, unreachable. Coverage is a coordination problem wearing a technical costume. The accent I replaced still survives in ~20 files. The token migration is real and incomplete. It moves file by file by design: there are no visual regression tests here, so bisectability is the only safety net, and batching would remove it. The 13px root is still in. It opens this case study, so it isn't hidden here, but it belongs on the unresolved list too, because it is unresolved. Every shared rule I wrote after it is in literal px so it can't be rescaled again, which is why the band rule says 44px and not h-11. The root itself is still 13px on the shipping branch. ### How I work ## Design against the real thing, not a picture of it No handoff, because there's nobody to hand off to. I design in Paper, build against the real components, and own the PR through review. Most of the decisions on this page are only visible* in the running product. A 13px root is invisible everywhere except the app, by which point it has rescaled a thousand things. Design files are where you decide what a thing should be. The product is where you find out what it is. --- # Building a content execution system for SMBs URL: https://anmolgupta.net/moflo-execution-system Client: MoFlo Role: Lead Product Designer Team: Product Designer (Me), CTO, Lead Developer, 2 Dev Interns Timeline: 5 Weeks Tools: Figma (Design, Make, Jam), Claude Code, Lottielab, v0, Fullstory Status: Shipped: Feb 2026 Summary: MoFlo could generate and schedule content, and users still weren't posting. The bottleneck wasn't generation quality, it was the cost of starting. I redesigned the product from a creation tool into a system that prepares work and asks for a decision. Every tool worked and nobody was posting. My first fix was the obvious one, and it barely moved. The thing that actually worked was giving up on getting users to start. Product tour of the MoFlo content execution system A queue the system prepares and the owner approves at a glance. — Numbers from the product's own analytics over the window I owned. The last one is the point: none of it came from new generation capability. - 60% weekly drop-off, the problem I was hired into, everyone opened it, nobody was posting (source: before) - 70% of system-generated drafts approved without a single edit (source: after) - 3× increase in multi-platform publishing per session (source: after) - 0 new generation capabilities added to get there: the model never changed (source: the whole argument) I was the only designer on this. That meant I also owned the part designers usually get to skip: writing down which behaviour we were betting on, running the experiment that could kill it, and telling the team when it did. ### The situation ## The product worked. The behaviour didn't. MoFlo is an AI content platform for small businesses: multi-platform publishing, AI generation, scheduling. All of it functioned. Users could generate captions and visuals and schedule across channels. They just didn't. % weekly drop-off, despite frequent dashboard visits posts per session before most users stopped % of users who acted on a performance insight Owners opened the dashboard, looked at the numbers, and left. The gap was behavioural, not technical. The platform before the redesign. — MoFlo platform overview ### The wrong guess ## I assumed the generation experience wasn't good enough My first hypothesis was the comfortable one: output quality. If the captions were better and the flow felt less uninspiring, engagement would follow. So I tested it before rebuilding anything. I simplified the input, surfaced suggested topic prompts, and cut friction out of the generation UI. The optimised generation flow, built to test the hypothesis rather than ship it. Drop-off improved slightly. Most users still stopped after one or two posts. The experiment was cheap and it was the most useful thing I did, because it ruled out the explanation everyone on the team preferred. ### What was actually happening ## Users weren't stuck on quality. They were stuck on starting. I ran interviews with active customers about how they approached content week to week. How do you decide what to post? When do you create content? What happens when you open the dashboard? What usually stops you from publishing? None of them ask about the product. If I'd asked what they thought of the generator I would have got opinions about the generator, and the answer was never going to be in there. Quote — SMB owner, customer interview: It takes too much time to create and schedule a post. I just want someone to do all of it for me. Quote — SMB owner, customer interview: I open the dashboard, look at the numbers, and then close it. It feels like I have to think too much before I can even start. The pattern: content creation was not part of a structured workflow. It was something they did when they had time or felt pressure. Every week started from zero. - Session 1 — Interview synthesis. Sorting what people said from what they did. - Session 2 — The same behaviour kept surfacing across different business types. - Session 3 — Grouping by the moment work stalled rather than by feature. - Session 4 — The through-line: deciding what to do next was the expensive step. They didn't lack tools. They lacked momentum. The friction was cognitive: the repeated mental effort of deciding, in a week already full of real work. ### The reframe ## Stop asking users to initiate If the expensive step is starting, the system should do the starting. That is the whole idea, and everything below is a consequence of it. Same job, both ways. The old flow spends four decisions before anything exists; the new one spends one after. - Existing flow - Decide to begin — Nothing happens until the user opens the dashboard and chooses to make something. There is no queue, no prompt, no starting point. - Decide what to post — A blank prompt. The user supplies the topic, the angle and the occasion before the AI can help at all. - Generate caption — The AI writes copy from whatever direction it was given. - Generate image — A second, separate generation step with its own inputs and its own wait. - Schedule — Pick a platform, pick a date, pick a time, confirm. - Redesigned flow - System detects a gap — Operational signals surface what needs attention: platform inactivity, a depleting content runway, an imbalance across channels. No user action involved. - Drafts are waiting — Content appears in an approval queue already tailored to the detected gap, so the session opens with work in it rather than a blank prompt. - Review in one place — How the post lands on Instagram, LinkedIn and X, all visible before approving. - Approve — Approval triggers optimised cadence. No manual date or time selection. Operational signals surfacing what needs attention, in real time. ### Getting there ## Three structures I built and rejected My earliest sketches split content by status so work was visible across stages, to make content feel like committed work rather than optional drafts. It didn't hold up. Users still had to decide what to do first. The system was organised but not decisive. Insights at the top, generated content below, calendar to the side. Metrics, action, schedule. This still required interpretation: read the signal, translate it into an action, navigate to drafts. It asked users to think in exactly the place they had already told me they stop. Early explorations put the calendar up front as a place to plan. But these users were not planners, they were reactors. So the calendar became a consequence instead: approval auto-fills it, hovering a scheduled day shows a lightweight preview rather than a heavy modal. Confirmation, not composition. - Pipeline sketches — Status columns. Organised, but still waiting on the user to choose. - KPI dashboard — Metrics first. Reads well, still asks for interpretation. - Calendar — Repositioned as a confirmation layer rather than a planning surface. - Final pipeline — Drafted, Flo Generated, Scheduled, Published. The second column is system-prepared work, not another bucket. The difference between the final pipeline and the first one is one column. "Flo Generated" is not a status the user moves things into. It is work the system did while they were gone. ### What shipped ## Review and approve, instead of plan and create Three pieces, in the order they run. I designed all three and prototyped each one before it went to the team, so what the developers got was a working thing to build against rather than a spec to interpret. Gap detection decides what to make. The engineering was straightforward; the design work was deciding what counts as a gap, and how loudly to say so. I scoped it to three signals, platform inactivity, a depleting content runway, an imbalance across channels, because those are the three an owner can act on in one click. Everything else I could have surfaced (engagement dips, follower deltas, best-time-to-post) fails that test: it's information, and information is what these users were already closing the dashboard on. Gap detection: inactivity, runway gaps and platform imbalance, surfaced without being asked. — Automatic gap detection One review, not three previews. A post that goes to Instagram, LinkedIn and X used to mean checking it three times in three places. I put all three renderings in one view, which sounds obvious and was the piece I had to argue for hardest: it costs real screen space, and the alternative: a platform switcher, costs a decision per platform, which is the exact currency the whole project was trying to stop spending. One unified review: how the post lands on each platform, before approval. Approval is the only decision left. Approving triggers an optimised cadence rather than opening a date picker. Users can still override the time, I kept that, because taking it away turns a system that helps into one that decides, but the default path has no scheduling step in it at all. Final execution system The pattern under all three: the system spends the decisions, the user spends the judgment. That split is why "AI does it for you" didn't have to mean handing over control, approval is a real veto, and the queue only exists to make using it cheap. ### Impact ## Consistency improved when initiative wasn't required Within one month: Weekly drop-off, one month after launch — The number the whole project was aimed at. Product analytics on the active base, one month post-launch, not a test cohort. % of system-generated drafts approved without edits × increase in multi-platform publishing new generation capabilities added to get there Users stopped planning content. They started maintaining it. The 70% figure is the one I'd interrogate if someone showed it to me. Approving without edits is only a good sign if the drafts were good; it looks identical to users rubber-stamping a queue to get through it. What separates the two here is the publishing number: rubber-stamped content doesn't get posted to three platforms. If those two had moved in opposite directions I'd have read the 70% as a failure, not a win. ### Reflection ## Activation is a design problem, not a model problem AI quality alone doesn't drive adoption. I spent the first stretch of this project improving output because that was the legible thing to improve, and the experiment that disproved it took two days. Behavioural architecture mattered more than feature depth. Reducing the cost of starting improved consistency without adding a single capability. --- # Designing a persona system for consistent AI content URL: https://anmolgupta.net/moflo-personas Client: MoFlo Role: Product Designer Team: Product Designer (Me), CTO, Lead Developer Timeline: 4 Weeks Tools: Figma (Design, Make, Jam), Claude Code, Lottielab, v0 Status: Shipped: Dec 2025 Summary: Brand voice drifted between sessions because users retyped it every time. I shipped a persona builder, watched it go unused, and rebuilt identity into the generation layer instead. I designed a system for people to describe their brand once. Half of them skipped it entirely and kept typing the same instructions every session. The second version stopped fighting that habit and absorbed it. Product tour of the MoFlo persona system Tone stopped being retyped into every prompt. — Measured against the sessions before the persona system shipped. - 2.3× increase in persona reuse across sessions (source: after) - 34% fewer manual tone instructions typed inside prompts (source: after) - 21% fewer regenerations caused by tone mismatch (source: after) I shipped v1, watched it fail, and made the case internally for rebuilding a feature that was already live, including writing up my own miss. The interesting half of this project is the second half, and I only got to it by being the person who said the first half hadn't worked. ### The situation ## AI is flexible. Brands are not. MoFlo could generate captions and visuals well. What it couldn't do was sound like the same company twice. - Users retyped tone instructions every session - Emotional tone differed across platforms - Manual edits to "fix voice" after every generation - Confusion about who the content was actually for The message that started this. Not an isolated complaint, but the first one that gave the problem language. — The user message that triggered the investigation The usage data had been hinting at it: high draft generation, low direct publish, frequent manual edits before scheduling. If every draft needs rewriting, the system isn't reducing cognitive effort. It's relocating it. The AI was producing content and users were still doing the thinking. ### What users actually meant by "brand" ## Nobody thinks in adjectives I expected people to describe voice the way a brand guideline does. They didn't. They thought in audience and intent. How they described voice — Nobody reached for adjectives. Every answer was about audience and intent, which is why a builder asking for tone, rules and stylistic preferences was asking the wrong question. We're premium but approachable. I don't know how to explain that to a model. I want it to sound like us, not like AI. Sometimes it's too salesy. That's not our brand. It forgets who we're talking to. After the builder shipped — Adoption didn't follow. People kept doing the manual thing because the manual thing was faster than remembering what the structured thing did. I just type what I want. I don't remember what this persona does. It's faster to just tell it again. What they could articulate: who they were speaking to, what they wanted to be known for, how they wanted to be perceived, and what they avoided saying. ### The version that failed ## I shipped a persona builder. Adoption didn't follow. The first system asked users to define brand identity explicitly, then apply it at generation time. Structured, deliberate, and more effort than most were willing to spend. - 1. Create — Define tone, audience, writing rules and stylistic preferences up front. - 2. Select — Choose the persona before generating. - 3. Apply — Rules applied behind the scenes, invisibly. Four weeks of usage data: Four weeks after the builder shipped. People kept doing the manual thing because it was faster. - Created at least one persona — 38 — 38% - Ever reused one — The number that mattered. Creation was a one-off; reuse was the point. — 17 — 17% - Skipped persona selection entirely — 50 — 50% Sixty percent still typed manual tone instructions into the prompt anyway, and regeneration for tone mismatch stayed high. The feature was structured, discoverable and shipped. It was also asking people to do setup work in exchange for a benefit they couldn't feel yet. That trade only works if the payoff is visible, and here it was invisible by design: the rules applied silently, so there was nothing to reward the investment. ### The rethink ## Stop asking users to build identity. Let them pick it and adjust. Rethinking the persona model The goal was never to eliminate prompting. It was to eliminate repetition. The difference reads small on paper. The reduction in cognitive load was not. - Before Users recreated their brand voice inside the prompt box every session. Tone adjustments lived in memory, not in the system, so every session started from nothing and the same instructions were retyped from scratch. - After Users selected an identity and adjusted it. The rules persisted between sessions, stayed visible during generation, and could be edited in place rather than rebuilt. ### What shipped ## Identity, moved into the generation layer Lightweight creation. Instead of constructing identity from scratch, persona creation starts from base archetypes and gets refined. Active visibility. The selected persona stays visible during generation. Users can see which rules are applied, edit them inline, and understand why output looks the way it does. Identity stopped being invisible and became explainable. Active persona visibility during generation The prompt-to-persona bridge. This is the piece I'd defend. When a user types a tone rule into the prompt ("make it less salesy", "more direct", "target investors"), the system offers to add it to the persona. Manual behaviour becomes structured data. Rather than fighting the habit of retyping, the system treats each retype as the user telling it something, and offers to remember. The habit that broke version one is what feeds version two. ### The hard part ## An offer that's wrong twice is an offer nobody reads again The bridge is easy to describe and was the piece I spent the longest on, because "detect a tone rule and offer to save it" hides three decisions that decide whether it works. The three calls inside one small prompt. Each one is a place I could have made the feature annoying enough to be ignored. - What counts — Only durable instructions get offered: a standing rule about voice, audience or things to avoid. One-off content direction (“mention the Tuesday special”) never does. - When to ask — After the generation lands, never before it. Interrupting on the way in adds a decision to the exact moment we were trying to make cheap, and the user hasn't seen the result they're judging yet. - What no means — Dismissing is per-rule and remembered. The same rule is never offered twice, whether or not the user ever adds it. I wrote the extraction rules as a design artefact, what qualifies, what doesn't, with examples on both sides, rather than handing engineering a confidence threshold to tune. The classification isn't mine. The definition of a false positive is, because a false positive here isn't a model error: it's the product telling someone it knows their brand better than they do. Version one asked for setup. Version two asks for confirmation. That is the entire difference, and it is worth more than any of the rules inside it. ### Impact ## Within one month × increase in persona reuse across sessions % fewer manual tone instructions inside prompts % fewer regenerations for tone mismatch Direct publish rate after the first draft rose 18%. The behaviour changed more than the numbers did: users stopped rewriting their identity every session and started selecting and refining it. One brief, written both ways. Reconstructed from the real outputs rather than exported. - No persona 🎉 BIG NEWS! We're now OPEN ON WEEKENDS! 🎉 That's right, you asked, we listened! Book your appointment today and experience the difference. Don't miss out! Link in bio 👇 #weekendvibes #bookNow #smallbusiness - Persona applied We're open Saturdays and Sundays starting this weekend. Same team, same hours you're used to on weekdays. If a weekday appointment has been the thing standing in your way, this is the fix. Booking link below. The difference isn't quality: the first one is competent marketing copy. It's that the first one could be any business, and the second one belongs to a specific business talking to people who already know it. That gap is what users were closing by hand, every session, with the same typed instruction. ### What I'd carry forward ## Structure has to absorb habit, not correct it AI features rarely fail for lack of capability. They fail when they misalign with what people already do. Version one asked users to change their behaviour first and rewarded them later. Version two watched the behaviour and built the structure around it. Reduce repetition, not control. And make the system's reasoning visible, because invisible correctness earns no trust. --- # Designing AI localization workflows for global media teams URL: https://anmolgupta.net/vosyn Client: Vosyn AI Role: Product Designer Team: Product Designer (Me), 2 Designers, UX Researcher, Product Manager, 2 Developers Timeline: Aug–Dec 2024 Tools: Figma (FigJam + Design), Balsamiq, Storybook, UserTesting Status: Handed off: 2024 Summary: Vosyn's AI could translate quickly. Reviewing what it produced meant moving between disconnected panels, so generation was fast and validation was slow. I redesigned the review workflow to keep generation, comparison and editing in one context. The translation took seconds. Checking it took minutes. I spent an internship on the unglamorous half of an AI product: not generating the output, but making it cheap to trust. Vosyn A localization workflow measured before and after. — SUS run with the same task set on both sides, so the two scores are comparable. - 64→82 System Usability Score, before and after (source: same task set) - 28% reduction in multi-step task time in localization workflows (source: after) - 5 AI-assisted workflow features prototyped and tested with users (source: shipped) This was an internship on a six-person product team, and the honest framing is that I owned two features inside a larger product rather than the product. The two below are the ones where the design call was mine end to end: the reframe that produced them, the patterns that came out of it, and the components that shipped into the team's Storybook. Under NDA, so this is written with redacted interfaces and no internal artefacts. Everything below is the reasoning and the shape of the work, which is the part that travels anyway. ### The situation ## Generation was fast. Verification was where the time went. Vosyn builds AI tools for translating and localizing content across languages. I joined as a product designer on the internal tooling used for AI-assisted translation and review: the surfaces the localization team lives in, not the ones customers see. I was brought in on a brief about the generation UI. The models were good and the workflow around them was not, and that turned out to be the whole project: reviewing a single translation meant moving between multiple panels, and reviewers struggled to compare original against translated content without losing their place. Reframing the brief from "improve the translation screen" to "the review is the product" was the first thing I did and the thing everything else came out of. The review interface as I can show it: redacted, and not much use to you. Everything below rebuilds the patterns instead. — Redacted localization review interface A consistent pattern in usability sessions: nobody complained about translation quality. They complained about finding out whether it was good. ### The reframe ## The unit of work isn't the translation, it's the check The unit of work isn't the translation. It's the check. Once I stopped treating the output as the product and started treating the review as the product, the design questions changed. Not "how do we present a translation" but "what does a reviewer need in view at the moment they decide it's fine". Two rules fell out of that, and they are the whole redesign. Under NDA I can't show the surfaces, but the shape of the work is the part that travels. Step through a single review, both ways. - Reviewing before - Open the segment — The reviewer lands on the translated output. The source sits on another surface. - Go find the original — Switch panels to read the source text that the translation is supposed to match. - Compare from memory — Judge the match with only one side in view at a time, going back and forth to check. - Leave to edit — Move to a different surface to make the correction that was just decided on. - Return and re-verify — Come back, find the place again, and confirm the edit landed. - Reviewing after - Open the segment — Source and translation are in view together. Comparison is simultaneous rather than sequential. - See what needs attention — The interface points at the places where the output is likely to need a human, instead of presenting every segment as equally suspect. - Edit in place — The correction happens on the same surface as the judgment that produced it. Getting there was unglamorous. I sat with reviewers working real files and timed where the minutes went, which is how the two rules above stopped being opinions: almost none of the time was spent reading translations, and almost all of it was spent re-establishing which two things were being compared. The UX researcher ran the formal sessions; I ran the workflow analysis and turned it into the patterns below, then prototyped and tested each one before it went near a sprint. The finding that changed my mind: reviewers were not slow readers. They were fast readers doing the same read three times, because each surface change meant re-finding their place. Nothing about that is a translation-quality problem, and no amount of model work would have touched it. ### What I designed · 01 ## Put the two texts in one view, and point at what needs a human The first pattern is the boring one that mattered most. Source and translation in a single view, so comparison is simultaneous rather than held in memory, and the rows likely to need a human are marked rather than leaving every row equally suspect. A rebuild of the review pattern, not the product. Click a flagged row to see why it was surfaced, or filter to just those. The flags are the design opinion here. A review tool that treats every segment as equally likely to be wrong spends the reviewer's attention evenly, which is the same as spending it badly. Routing attention is the feature. Which is why the flag categories are the part I'd defend. They are not confidence scores; a low-confidence score tells a reviewer that the model is unsure, which they can neither check nor act on. Each flag names a failure mode a human can adjudicate in one read: an idiom carried over literally, a register that drifted formal, a glossary term that resolved inconsistently. And each one shows its reasoning, because a flag a reviewer can disagree with is a flag they will keep using, while an opaque one gets switched off the first week it is wrong. The limit I'd name up front: this only routes attention toward failure modes we thought to define. A model error in a category nobody anticipated looks exactly like a clean segment, and a reviewer trusting the flags is now skimming it faster than before. Flagging makes good reviewers faster; it does not make an unreviewed file safe, and I'd want that written into how the feature is introduced, not just into how it works. ### What I designed · 02 ## Answers that name where they came from The second piece was the contextual inquiry system: asking questions against a selected stretch of video and getting responses grounded in that segment rather than the whole file. The design problem is trust. An assistant that answers from everywhere gives you nothing to check. Scoping the question to a segment, and having the answer say which segment it used, turns a claim into something verifiable. A rebuild of the grounding pattern, not the product. Select a segment, then ask: the answer cites the range it was drawn from. 00:00–00:42 00:42–01:35 01:35–02:10 Same principle as the flags: the interface's job is to make checking cheap. A grounded answer can be disagreed with. An ungrounded one can only be believed or ignored. ### Outcomes ## Measured after usability testing System Usability Score, usability testing — SUS moved from the low band into the high band. Measured on a test cohort during the internship, not production telemetry. System Usability Score before, rising to 82 after % reduction in multi-step task time in localization workflows AI-assisted workflow features prototyped and tested Teams validated machine translations faster, and the friction that had been sitting after generation moved out of the way. Worth stating plainly: these come from usability testing during the internship, not from production telemetry after launch. The SUS movement is real and measured, on a test cohort rather than the full user base. ### Reflection ## Generating is half the product The lesson I took from this one, and have used since: for AI products, producing the output is rarely the bottleneck. Evaluating it is. Users need efficient ways to check, correct and trust what the model gives them, and that surface gets far less design attention than the generation flow does. Small changes to workflow structure moved the needle more than any new feature would have. Reducing the cost of verification is a product strategy, not a polish task. ---