Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Scalable, evolving training environments for computer-use agents — deepening the task distribution agents learn from.
Computer Use EnvironmentsAs a technical program leader, I’ve shipped 9 frontier models and agentic systems at Microsoft AI Superintelligence on an 8-week release cadence. I run release, evaluation, and safety across reasoning models and computer-use agents — which mostly means getting research, engineering, Responsible AI, legal, and go-to-market pointed at the same launch date.
Before Microsoft I spent three years at Google, leading generative AI product strategy for Shopping. I hold an M.S. in Human-Computer Interaction from Georgia Tech and a B.Tech from NIT Karnataka.
A model isn’t shipped when it’s trained. It’s shipped when evaluation, safety, legal and go-to-market all agree it’s ready — on the same day.
Outside work I read a lot, sing, paint, and bake — currently in a lavender mint cake phase.
A short version. The projects below are where the detail lives.
I own the product portfolio for Microsoft's small language model and agentic AI programs, and built the release pipeline every model now ships through.
Owned generative AI product strategy for Shopping, from the Search Generative Experience to Virtual Try-On.
Strategy and problem-solving across industries, before moving into tech.
Scalable, evolving training environments for computer-use agents — deepening the task distribution agents learn from.
Computer Use EnvironmentsThe Fara1.5 portfolio — a three-tier 4B/9B/27B family whose 27B model reached 72% on Online-Mind2Web, surpassing OpenAI Operator and Gemini 2.5 Computer Use.
Agentic AI Computer UseA 7-billion parameter computer-use model that beats GPT-4o at roughly 1/12th the inference cost, with a Critical Points framework requiring human approval before irreversible actions.
Agentic AI Computer UseThe 14B reasoning model that matches DeepSeek-R1 (671B) on AIME 2025, establishing Microsoft's position in efficient reasoning.
Reasoning Language ModelsInvestigation of inference-time compute scaling methods across reasoning tasks, exploring how additional compute at inference improves performance.
Scaling ReasoningA framework for synthetic data creation using agentic flows — its 25M instruction pairs became the training foundation for every subsequent Phi model.
Synthetic Data AgentsThe work I’m closest to, most recent first. Each tile says what my role actually was.
The agentic product surfacing this work — market research through product-market fit, and GTM through launch.
Set the three-tier 4B/9B/27B strategy and the data bet behind it. The 27B model reached 72% on Online-Mind2Web.
Microsoft’s bet on efficient computer-use agents — beating GPT-4o at roughly a twelfth of the inference cost.
Took the Phi reasoning program from idea to shipped models, landing a 14B that matches DeepSeek-R1 at 671B on AIME.
One readiness path every model now ships through, replacing per-launch scramble with a repeatable process.
The safety pattern for computer-use agents: a human approves before the agent does anything irreversible.
Chose synthetic data over human annotation — 25M instruction pairs that became the training base for every Phi model since.
Shopping’s top AI initiative — 7% lift in user satisfaction and 3% lift in impressions across Google Search.
Shipped Virtual Try-On as part of owning Shopping’s generative AI roadmap.
Also: won the Executive Challenge at Microsoft’s Global Hackathon 2025, and speak and review at MLADS.
Three things I'm responsible for. Hover or scroll to move between them.
Release management, launch readiness reviews, dependency tracking, and Responsible AI review — the machinery that gets a model from ready to shipped.
Reasoning and multimodal models, agentic systems, and the evaluation and red-teaming work that decides whether they are safe to release.
The tooling underneath — multi-agent frameworks, serving, and the platforms models actually ship on.
Best if it’s about shipping frontier models — evaluation, release process, computer-use agents, or how safety review actually works in practice. I read everything.
Co-Intelligence: Living and Working with AI
Ethan Mollick
Reading, singing, painting, baking
Currently a lavender mint cake phase
San Francisco Bay Area
Previously Vancouver, Singapore, and Muscat