TechnicalNews

Microsoft Fara1.5 27B NEW Browser Automation Model is WILD!

Julian Goldie SEO

Microsoft released Phi-3.5, a family of three computer use models (4B, 9B, 27B) that automate browser tasks through vision-based clicking rather than HTML parsing. These open-weight models significantly outperform larger closed-source alternatives like OpenAI's Operator and Google's Gemini 2.0 on web automation benchmarks.

Summary

Microsoft Research has released Phi-3.5, a family of three browser automation models built on Qwen 3.5 with weights available under MIT license on Hugging Face. Unlike traditional chat models that read and write text, or web agents that parse HTML code, Phi-3.5 uses vision-only technology to observe browser screenshots, reason about what it sees, and predict pixel coordinates to click—mimicking how humans interact with webpages. This approach proves superior when pages use custom widgets or unconventional code structures that confuse traditional agents.

The three model sizes (4B, 9B, 27B) offer different performance and resource trade-offs. The 27B variant holds a 262K token context window, maintaining the three most recent screenshots while storing the rest as plain text. On the Online Mine 2 Web benchmark (300 tasks across 136 websites), the 27B scores 72.3%, outperforming OpenAI's Operator (58.3%), Google's Gemini 2.5 Computer Use (57.3%), and University of Toronto's Navigator (64.7%). Performance drops on longer, cross-site tasks in the Web Tail Bench benchmark, with the 27B achieving 40.2% on outcome success—a limitation Microsoft acknowledges transparently.

The models incorporate safety features through a trained observation of critical points: missing information (stopping to ask for required details rather than fabricating them), unclear tasks, and irreversible actions like form submission or account login (requiring prior authorization). Microsoft strongly recommends running Phi-3.5 only within sandboxed environments like their harness or Magentic Light, which provide containerized browsers with no file access, domain allow-lists, full action logging, and instant pause capabilities. The speaker emphasizes that screen size (1440x900 for optimal accuracy), step limits (starting with 10-20 for testing), and specific outcome descriptions significantly impact performance. The 4B model is recommended for local deployment as it still outperforms OpenAI's Operator while consuming minimal resources.

Key Insights

  • Phi-3.5 uses vision-only approach to read browser pages like humans do and predict pixel coordinates to click, rather than parsing HTML code, making it more robust when pages use custom widgets or unconventional structures
  • The 27B Phi-3.5 model outperforms significantly larger closed-source models—scoring 72.3% on Online Mine 2 Web benchmark compared to OpenAI Operator's 58.3% and Gemini 2.5's 57.3%
  • Phi-3.5 is trained with three types of critical points where it stops and asks for authorization: missing information scenarios, unclear task specifications, and any actions that cannot be undone like form submission or account login
  • Performance on long multi-step cross-site tasks remains significantly lower than on single-site benchmarks, with Web Tail Bench scores of 40.2% for the 27B, indicating this remains a hard problem
  • Screen resolution matching (1440x900 where models were trained) noticeably improves clicking accuracy, and the 4B model still outperforms OpenAI's Operator while being deployable locally

Topics

Phi-3.5 browser automation modelsVision-based computer use vs HTML parsingBenchmark performance comparisonsSafety mechanisms and critical pointsLocal deployment and hardware requirementsWeb automation workflow examplesModel sizing and performance trade-offs

Transcript

[0:00] Microsoft Fara 1.5 27B new browser automation model is wild. What if the AI you use every day is only doing half the job? Everyone is chasing chat models. Microsoft just dropped something else. This one doesn't talk back. It clicks and the smallest version runs on your own machine. So why is almost nobody talking about it? I'm the digital avatar of Julian Goldie and I help you learn AI tools and actually use them in your work. Stick with me because in a minute I'll walk you through the exact demo and I'll show you the one setting most [0:30] people leave switched off on their first try. Let's start with what Fara actually is. Fara 1.5…

Full transcript available for MurmurCast members

Sign Up to Access

More from Julian Goldie SEO

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.