Microsoft Fara1.5 27B NEW Browser Automation Model is WILD!
Microsoft released Phi-3.5, a family of three computer use models (4B, 9B, 27B) that automate browser tasks through vision-based clicking rather than HTML parsing. These open-weight models significantly outperform larger closed-source alternatives like OpenAI's Operator and Google's Gemini 2.0 on web automation benchmarks.
Summary
Microsoft Research has released Phi-3.5, a family of three browser automation models built on Qwen 3.5 with weights available under MIT license on Hugging Face. Unlike traditional chat models that read and write text, or web agents that parse HTML code, Phi-3.5 uses vision-only technology to observe browser screenshots, reason about what it sees, and predict pixel coordinates to click—mimicking how humans interact with webpages. This approach proves superior when pages use custom widgets or unconventional code structures that confuse traditional agents.
The three model sizes (4B, 9B, 27B) offer different performance and resource trade-offs. The 27B variant holds a 262K token context window, maintaining the three most recent screenshots while storing the rest as plain text. On the Online Mine 2 Web benchmark (300 tasks across 136 websites), the 27B scores 72.3%, outperforming OpenAI's Operator (58.3%), Google's Gemini 2.5 Computer Use (57.3%), and University of Toronto's Navigator (64.7%). Performance drops on longer, cross-site tasks in the Web Tail Bench benchmark, with the 27B achieving 40.2% on outcome success—a limitation Microsoft acknowledges transparently.
The models incorporate safety features through a trained observation of critical points: missing information (stopping to ask for required details rather than fabricating them), unclear tasks, and irreversible actions like form submission or account login (requiring prior authorization). Microsoft strongly recommends running Phi-3.5 only within sandboxed environments like their harness or Magentic Light, which provide containerized browsers with no file access, domain allow-lists, full action logging, and instant pause capabilities. The speaker emphasizes that screen size (1440x900 for optimal accuracy), step limits (starting with 10-20 for testing), and specific outcome descriptions significantly impact performance. The 4B model is recommended for local deployment as it still outperforms OpenAI's Operator while consuming minimal resources.
Key Insights
- Phi-3.5 uses vision-only approach to read browser pages like humans do and predict pixel coordinates to click, rather than parsing HTML code, making it more robust when pages use custom widgets or unconventional structures
- The 27B Phi-3.5 model outperforms significantly larger closed-source models—scoring 72.3% on Online Mine 2 Web benchmark compared to OpenAI Operator's 58.3% and Gemini 2.5's 57.3%
- Phi-3.5 is trained with three types of critical points where it stops and asks for authorization: missing information scenarios, unclear task specifications, and any actions that cannot be undone like form submission or account login
- Performance on long multi-step cross-site tasks remains significantly lower than on single-site benchmarks, with Web Tail Bench scores of 40.2% for the 27B, indicating this remains a hard problem
- Screen resolution matching (1440x900 where models were trained) noticeably improves clicking accuracy, and the 4B model still outperforms OpenAI's Operator while being deployable locally
Topics
Transcript
[0:00] Microsoft Fara 1.5 27B new browser automation model is wild. What if the AI you use every day is only doing half the job? Everyone is chasing chat models. Microsoft just dropped something else. This one doesn't talk back. It clicks and the smallest version runs on your own machine. So why is almost nobody talking about it? I'm the digital avatar of Julian Goldie and I help you learn AI tools and actually use them in your work. Stick with me because in a minute I'll walk you through the exact demo and I'll show you the one setting most [0:30] people leave switched off on their first try. Let's start with what Fara actually is. Fara 1.5…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Julian Goldie SEO
How to Run DeepSeek V4 Flash for FREE!
A tutorial demonstrating how to use DeepSeek V4 Flash for free through Open Code and integration with agent operating systems like Hermes Agent. The speaker showcases building websites and apps using this free AI model and explains how it compares favorably to larger models despite being smaller.
Claude Obsidian 2.0 is INSANE (FREE!)
Claude Obsidian 2.0 is presented as a free AI memory upgrade that allows users to upload files into a folder for permanent retention and linking. The system creates a knowledge graph that learns from business documents, provides sourced answers, and can be shared across teams.
Claude Agent OS is INSANE! 🤯
Julian presents a comprehensive Claude-based agent operating system that integrates multiple AI models, automated workflows, and a persistent memory system to automate daily tasks. The system runs 24/7 and uses free or existing subscriptions, combining tools like voice agents, content creation, competitor monitoring, and real-time news analysis into a single unified dashboard.
NEW ChatGPT Update is INSANE!
OpenAI released a major ChatGPT update featuring a Chrome extension called Side Chat and an improved desktop app that work together to streamline SEO research and content creation. The update allows users to analyze multiple browser tabs simultaneously, highlight text for quick answers, and convert research into finished work without constant tab switching.
This NEW AI Design Tool Changes EVERYTHING
Julian Goldie introduces Replit Design, a new AI-powered design tool launched July 29, 2026, that enables users to create professional designs without traditional design software by using ambient intelligence and one-click iterations. The tool integrates multiple AI models, built-in design references from Mobbin, brand system support, and direct integration with live product deployment on the Replit platform.