Taking an agentic-first developer cloud from 0 → 1

A research program that carried an unreleased platform from "should we build this?" to a launch-ready experience — across generative research, a concept value test, usability testing, and a naming study.

  • Role: Research lead (end-to-end)

  • Methods: Generative · concept value test · usability · naming

  • Context: Unreleased developer cloud platform, Microsoft

What this work demonstrates

Research method design · Generative + evaluative range · Qual + quant synthesis · 0→1 product judgment · Stakeholder influence · Data visualization

Context and Stakes

Microsoft was exploring an unreleased concept: a cloud platform that strips the assembly out of shipping software — connect a repository, deploy through natural-language, agent-driven workflows, with the production essentials wired in and a path to graduate into the company's full hyperscale cloud when you outgrow the defaults.

The bet was real, and so was the ambiguity. The concept lived in a crowded middle: lightweight "neoclouds" had already won developers on speed and simplicity, while hyperscalers owned scale and control. Nobody knew whether developers wanted something in between, whether they'd trust an AI agent to touch their infrastructure, whether the first build was usable — or even what to call it. There was no product yet, just a thesis and a lot of open questions.

My role

I led the research from framing to influence, partnering with one other researcher. Across the program I owned the business questions, the method strategy — choosing the right instrument for each stage — study design, moderation, synthesis of both qualitative and ranking data, and the readouts that carried findings into product, design, and marketing.

The program — four studies, four decisions

Phase 1 · Understand — Generative research

opened with two rounds of foundational interviews, split across segments: enterprise and regulated pro-developers first, then startup and hobbyist builders. The method was semi-structured interviewing paired with journey mapping and explicit, pre-registered hypotheses I set out to validate or kill.

Developers trust AI for code — skeletons, tests, snippets — but not for autonomous production deployment. Deployment was the trust cliff: the moment a wrong move costs money or breaks something they'd have to explain to their team. Human-in-the-loop control was non-negotiable, cost and security were hard gates, and developers lived in fragmented toolchains they wanted stitched together. Startup builders tolerated far more abstraction than enterprise developers, who needed transparency, approvals, and an escape hatch to the underlying cloud.

This phase defined the opportunity and its guardrails: an opinionated, time-saving layer that keeps a human in the driver's seat, with visible plans and a way down to the raw cloud — plus the hypotheses the concept itself would have to satisfy.

Phase 2 · Validate — Concept value test

With a concept on the table but no prototype, I had to test value — and the usual failure mode of concept research is that an exciting pitch produces aspirational agreement that evaporates at adoption. So I designed a forced-tradeoff, side-by-side concept value test: participants ranked the benefits and limitations of the cloud they use today; I normalized those into shared dimensions; introduced the concept strictly as a neutral comparator; had them judge better / worse / same on each dimension; then forced a choice and named the opportunity cost. Behavioral proxies ("which would you pick for your next project?") replaced unreliable stated switching intent.

Insight 1 — The biggest draw and the biggest blocker were the same thing. Developers wanted agent-driven automation, but only with visible control — security and human-in-the-loop approval were the gate, not an edge case.

Insight 2 — The "middle" is real, but only as opinionated defaults with a visible escape hatch. Developers valued a position between lightweight simplicity and hyperscale control, but not as a black box.

Insight 3 — Cost predictability was a gating factor, not a feature. Fear of surprise bills from opaque, auto-scaling pricing was a recurring adoption brake; transparent, predictable cost was a precondition for trust.

Phase 3 · Refine — Usability testing

Once there was a first working build, I ran task-based usability sessions — find and deploy a template, explore what was created, make a change and redeploy, choose a workflow across portal, IDE, CLI, and AI. I evaluated against usability heuristics, rated issue severity, and captured perceived-ease ratings.

The headline wrote itself: the first deployment is easy; everything after that needs work. The fast start genuinely delighted — but developers couldn't tell what they'd just created, a successful deploy could look like a failure without clear confirmation, and the paths to logs and recovery were hard to find. Speed without comprehension is a false success.

This phase pushed the team to fix post-deploy comprehension and system-status feedback — the unglamorous part — ahead of adding surface area.

Phase 4 · Position — Naming study

Finally, the name. I ran a participatory naming study — participants generated and evaluated candidate names, which I assessed against clarity, memorability, and credibility. The finding was clear and a little inconvenient: for developer tools, clarity beats creativity. Function-signaling names won; brand-heavy, playful names created friction and doubt about what the product even did. Trust came from the company behind the tool far more than from the name itself. I brought the uncomfortable version forward — the concept's own working name tested as a source of friction, and I said so rather than burying it.

Impact — what changed because of the work

Taken together, the program shaped the product from thesis to launch-readiness.

  • Set the product thesis and its guardrails — generative research defined the opportunity and named the trust, cost, and security constraints any solution had to respect.

  • Validated the concept and its positioning — the concept value test gave leadership a defensible read on demand and a clear position for the crowded middle.

  • Made trust a design pillar — human-in-the-loop control and security transparency moved from afterthought to positioning pillar; cost became an inline decision companion.

  • Reprioritized the roadmap — usability findings pushed the team to fix post-deploy comprehension and status feedback before scaling features.

  • Informed the name and brand — the naming study fed a more defensible, clarity-first direction.

  • Traveled cross-functionally — findings moved into product, design, and product-marketing decisions, not just a research deck.

Reflection — what I'd do differently

Carry a consistent participant pool across phases. Segment earlier and, where possible, run a lightweight longitudinal panel so insights compound instead of resetting each study.

  • Bring behavioral validation into the concept phase sooner. Even a thin prototype task would let me check stated tradeoffs against observed behavior earlier.

  • Standardize the live synthesis step. Normalizing priorities into shared dimensions in real time is powerful but introduces moderator influence — I'd systematize and validate it between sessions.

  • Pre-commit the analysis for small samples. I'd fix aggregation methods in advance and add a lightweight quantitative layer to size qualitative signals.

  • Tie naming back to positioning explicitly so brand and value proposition are argued from one evidence base.