Where people stop trusting automation — UX Research Case Study

UX Research · Case Study

Where people stop trusting automation

Taking an unreleased developer cloud from “should we build this?” to a shipping plan

Four studies over three months, each sized to the decision in front of the team: whether to build it, what to build first, whether the first version worked, and what to call it.

Role: Research lead, end to end Methods: Foundational interviews · concept testing · usability testing · naming research Context: Unreleased cloud platform, Microsoft
Concept through launch planning Exploratory and evaluative Shaped product, design and marketing

What this work demonstrates

Research method design Exploratory and evaluative range Qualitative and quantitative synthesis Judgment on a product with no users yet Framework building Reusable research process

The short version

What happened, in five sentences

A team wanted to build a service where an AI agent does the work of getting an application from finished code to actually running. Nobody knew whether people wanted it, whether they would trust it, or what to call it. I ran four studies over three months, each one sized to the decision the team had to make next.

The finding that mattered most was about trust: people were glad to let the AI write code and refused to let it deploy that code without them. Across three studies I pulled that into four dimensions of trust, each one traceable from a specific piece of evidence to a specific product decision. Almost everything the team decided, including what to build first and who the product was for, came back to one of those four.

Context and stakes

A last-mile problem at the moment of choice

Software development was shifting fast. Applications were starting to build AI in as a standard component, and a much larger wave of people who write software without being professional developers was entering the market.

180M+developers on GitHub, and growing fast
4:1people who build software without being professional developers, projected to outnumber those who are

That group picks whichever cloud lets them ship fastest, and the Azure developer services team had a problem at the last mile: the gap between working code and a running production application. Azure was powerful but complicated, and simpler competitors were winning this audience with a frictionless push-code-get-app experience the enterprise offering could not match. The bet was a simplified, AI-native platform that could win people at the moment they choose, on brand-new projects, and graduate them into the full cloud as they grew.

The concept stripped the assembly out of shipping software. You connect a code repository and deploy through agent-driven workflows, meaning the AI carries out the steps rather than just suggesting them, with the production essentials wired in and a path up to the full-scale cloud when you outgrow the defaults.

What the bet lacked was evidence. Did anyone want this, would they trust an agent with their infrastructure, was the first build usable, and what should it be called?

I owned the research that answered those questions in sequence, in a four-phase program with each phase sized to the decision in front of the team.

My role

Led the research from framing to influence

I led the research from framing to influence, partnering with one other researcher. Across the program I owned the business questions, the method strategy of choosing the right instrument for each stage, study design, moderation, synthesis of both open-ended and ranking data, and the readouts that carried findings into product, design and marketing decisions.

I also used the program to onboard a contract researcher onto the product team: reviewing materials, modeling moderation, and coaching them through leading sessions, while staying accountable for program direction and the final recommendations.

The program

Four studies, four decisions

Four research phases in sequence Phase one, understand: did developers want this at all, two rounds of foundational interviews. Phase two, validate: would they actually switch, a forced-tradeoff concept test. Phase three, refine: was the first build usable, task-based usability sessions. Phase four, position: what do we call it, a participatory naming study. Three months, with the product unreleased throughout. 1 UNDERSTAND Did developers want this at all? 2 rounds of foundational interviews 2 VALIDATE Would they actually switch? Forced-tradeoff concept test 3 REFINE Was the first build usable? Task-based usability sessions 4 POSITION What do we call it? Participatory naming study 3 months · product unreleased throughout
Each study was sized to the decision in front of the team, not to a fixed research calendar.
Understand · Foundational research

Did developers want this at all?

I opened with two rounds of foundational interviews, split across groups: professional developers in enterprise and regulated settings first, then people building at startups and on their own. The method was semi-structured interviewing paired with journey mapping and explicit hypotheses I set out in advance to validate or kill.

What came out of it was not a list of feature requests. It was a boundary.

Trust is positional, not general

Developers trusted the agent to write code and refused to let it deploy. Trust didn’t fade gradually, it ended at the point where a mistake became expensive and hard to undo. Human review at deployment was non-negotiable, cost and security were hard gates rather than preferences, and less experienced builders tolerated far more automation than professional developers, who wanted transparency, approval steps, and a way back down to the system underneath.

The shape of this finding isn’t specific to software. Automation gets accepted right up to the point where a mistake becomes expensive and hard to attribute, and that boundary is where the research has to concentrate.

This phase defined the opportunity and its guardrails at the same time: an opinionated layer that saves real time while keeping a person in the driver’s seat, with visible plans and a route down to the raw cloud.

Validate · Concept testing

Would they actually switch?

There was a concept on the table but no prototype, so I had to test value rather than usability. The usual failure mode here is that an exciting pitch produces agreement that evaporates at adoption.

I designed a forced-tradeoff concept test, meaning participants compare the proposed product against what they already use and are made to choose between them rather than rate both favourably.

Five-step forced-tradeoff concept test Step one, elicit: rank what you value about the tool you use today. Step two, normalize: turn those into shared dimensions. Step three, introduce: show the concept as a neutral comparator. Step four, compare: judge better, worse or same on each dimension. Step five, choose: pick one and name what you give up. 1 Elicit Rank what you value about the tool you use now 2 Normalize Turn those into shared dimensions 3 Introduce Show the concept as a neutral comparator 4 Compare Better, worse or same on every dimension 5 Choose Pick one and name what you give up The last step is the one that works. Enthusiasm is free; naming the cost isn’t.
Built to surface what people would actually give up, rather than what sounds appealing described out loud.

The biggest draw and the biggest blocker were the same thing

Developers wanted agent-driven automation, but only with visible control. Security and human approval were the gate, not an edge case. The capability and the anxiety were two sides of one feature.

Cost predictability was a gate, not a feature

Five of eight participants raised unexpected or unpredictable billing without being asked, including one who had received an accidental $35,000 bill. Fear of a surprise charge from opaque, auto-scaling pricing was a live adoption brake. Cost ranked first among the six migration motivators I classified.

Five of eight said they would consider migrating, but their current platforms were cheap enough at their usage level that cost decided it. Alongside that, a segment finding narrowed the target: developers in enterprise settings needed more control and customization than a lightweight platform could offer and generally needed a full-scale cloud, which pointed the initial audience at startup and independent developers instead.

Positioning map of speed against control Two axes: speed to a running app on the vertical, control and customization on the horizontal. Lightweight platforms sit high on speed and low on control. Full-scale cloud sits low on speed and high on control. The gap developers described wanting is high on both, in the upper right. SPEED TO A RUNNING APP CONTROL AND CUSTOMIZATION low high Lightweight platforms fast to ship, little room to adapt Full-scale cloud endless control, slow to first result The gap opinionated defaults, with a visible way down to the system underneath
Developers wanted the middle, but only as sensible defaults they could see through and step around, never as a black box.
Refine · Usability testing

Was the first build usable?

Once there was a working build, I ran task-based sessions: find and deploy a template, explore what had been created, make a change and redeploy, and choose a path across the web portal, the code editor, the command line, and the AI. I evaluated against usability heuristics, rated issue severity, and captured how easy people found each task.

Usability results before and after the first deployment The first deployment succeeded for all eight participants and was rated easy. After changing their own code, six of eight could not redeploy without assistance. Nineteen issues were logged in total, three rated high severity. THE FIRST DEPLOYMENT 8 of 8 deployed a template successfully, and said so This part genuinely delighted people. EVERYTHING AFTER IT 6 of 8 could not redeploy without help after changing code A successful deploy could look like a failure. 19 issues logged · 3 rated high severity · roadmap reordered toward what happens after the first success
Speed without comprehension reads as a win in the metrics and a failure at the desk.

The pattern was consistent: the first deployment was easy, and everything after it was not. The fast start genuinely delighted people, but they could not tell what they had just created, a successful deploy could look like a failure without clear confirmation, and the routes to logs and recovery were hard to find.

How I decided what to fix first

Findings pass a five-check quality rubric, then get scored on frequency, impact, persistence and scope, with impact weighted above frequency. Every rating carries written rationale and sample evidence, so the ranking can be argued with rather than taken on faith.

I review and finalize each one and add a feasibility pass, which sometimes deprioritizes a real finding because the cause sits outside the product. Building the model mattered as much as the study: it gave a team with no research history a transparent way to contest a priority instead of just accepting or ignoring it.

Position · Naming and go-to-market

What do we call it, and how do we talk about it?

The final phase looked ahead to launch. I ran a participatory naming study where participants generated and evaluated candidate names, which I then assessed against clarity, memorability and credibility.

The finding was clear and a little inconvenient: for developer tools, clarity beats creativity. Names that signalled function won. Brand-heavy, playful names created friction and doubt about what the product even did. Trust came from the company behind the tool far more than from the name itself.

I brought the uncomfortable version forward. The concept’s own working name tested as a source of friction, and I said so rather than burying it.

The study set the marketing foundation, not just the name

How developers described the product back to us, and what earned or lost their trust, became direct input to how it would be marketed: lead with what it does and how fast, let the established brand carry credibility rather than inventing a new one, and treat human review and predictable cost as first-class selling points, because those were the terms on which developers said they would adopt.

Synthesis

The trust framework

“Developers don’t trust the agent” isn’t actionable. Across three studies I pulled the trust findings into four dimensions, each one traceable from a specific piece of evidence to a specific product decision.

The traceability is the point. A framework that can’t be walked backward to the evidence is an opinion, and a framework that can’t be walked forward to a decision is decoration. This one is the spine of what the team built.

How trust dimensions shaped the product

Four dimensions identified across three studies with 32 developers

Developers were uncomfortable with an agent making higher-risk changes without review.

ControlCustomers decide what the agent can do without approval.

Agent shows its plan; approval points added before higher-risk actions.

Participants could deploy successfully but were unsure what the system had created.

TransparencyCustomers understand what happened after an action.

Clearer deployment summaries, system-status feedback, and contextual guidance.

Cost ranked first among the six migration motivators.

PredictabilityCustomers understand the likely financial impact.

Cost visibility and predictable defaults set as core product requirements.

Six of eight participants could not redeploy without assistance after changing code.

ComprehensionCustomers hold a mental model of how code moves from Git to the cloud.

GitHub integration prioritized, deployment workflow explained, and target customer refined by source-control knowledge.

Why four dimensions rather than one score

Trust isn’t a quantity people have more or less of. It’s a set of conditions, and different products fail different ones. Splitting it apart meant a team could ask which dimension a given feature was serving, and notice when a change improved one while quietly damaging another.

None of the four is specific to cloud infrastructure. Any system that acts on someone’s behalf has to answer all four, and the ones it can’t answer are where adoption stalls.

Impact

What changed as a result

Four launch-critical decisions were made on the strength of this research, before a line of the product was public.

The research wasn’t running alongside the product; it was the path the product took. Each phase answered the decision that unlocked the next: whether to build it, what to build first, what to fix, and how to talk about it.

  • Set the product thesis and its guardrails. Foundational research defined the opportunity and named the trust, cost and security constraints any solution had to respect.
  • Made trust a design pillar, in four specific parts. Control, transparency, predictability and comprehension each turned into concrete product requirements rather than a general commitment to being trustworthy.
  • Redefined the target customer. A segment finding narrowed the initial audience from enterprise to startup and independent developers, because source-control fluency turned out to gate whether the concept made sense at all.
  • Changed the shape of the product. A standalone portal instead of the broader Azure portal, with initial scope prioritizing repository integration, templates, contextual guidance and approval controls.
  • Reordered the roadmap. Usability findings pushed the team to fix comprehension and status feedback after deployment before scaling features.
  • Informed the name and the go-to-market. Naming research fed a clarity-first name direction and the messaging pillars for the launch narrative.
  • Left behind reusable process. The trust framework, the severity model and the workshop-style handoff outlived the individual studies.
  • Traveled cross-functionally. Findings moved into product, design and product marketing decisions, not just a research deck.

Reflection

What I’d do differently

  • Carry a consistent participant pool across phases. Segment earlier and, where possible, run a lightweight ongoing panel so insights compound instead of resetting each study.
  • Bring behavioral validation into the concept phase sooner. Even a thin prototype task would let me check stated tradeoffs against observed behavior earlier.
  • Standardize the live synthesis step. Normalizing priorities into shared dimensions in real time is powerful, but it introduces moderator influence. I’d systematize and validate it between sessions.
  • Pre-commit the analysis for small samples. I’d fix aggregation methods in advance and add a lightweight quantitative layer to size the qualitative signals.
  • Test the trust framework prospectively. It was built from findings after the fact. The stronger version would use it to generate predictions and then check them.

About this case study. This work has been sanitized for public sharing. Internal code names, parent-service identifiers, unreleased details and exact internal figures have been removed or abstracted. The product had not launched at the time this program concluded.

Taryn Bipat · UX Research · Developer cloud research program, Microsoft