Agentic Vision Is Here… AI Now Sees, Thinks, and Runs Work Without You (Crazy..)
The video introduces 'agentic vision,' a capability from Gemini that allows AI to observe images, screens, and dashboards, understand them like a human, and take action autonomously. The speaker argues this represents a new frontier in AI that most people are unprepared for. The video is largely promotional, directing viewers to comment for a free training course.
Summary
The video opens with the host introducing a technology called 'agentic vision,' which he attributes to Gemini (referencing 'Gemini 3'). He describes it as a capability that allows AI to visually perceive and interpret images, screens, dashboards, and other visual environments in a way that mirrors human understanding.
The host emphasizes that agentic vision goes beyond passive observation — the AI can also plan, make decisions, and take actions based on what it sees, all without requiring constant human supervision or instruction. He frames this as a significant departure from traditional AI interaction models, where users craft prompts to direct the AI.
The speaker expresses a sense of urgency and excitement, claiming that most people are unprepared for the implications of AI that can autonomously observe and act within real-world and digital environments. He positions this as a 'new frontier' of AI capability sourced directly from Google.
The majority of the short transcript is devoted to promotional calls-to-action, asking viewers to comment 'guide' or 'start' in order to receive a free four-part fast track training program related to the topic.
Key Insights
- The speaker claims that 'agentic vision' allows AI to not just see but understand visual environments — including images, screens, and dashboards — in the same way a human would, enabling it to decipher, plan, and make decisions.
- The speaker argues that agentic vision signals a shift away from prompt engineering as the primary mode of AI interaction, since the AI can now observe and act on its own without being 'baby sat.'
- The speaker claims this capability extends to real-world digital environments, images, and video — not just static inputs — and that the AI can take action based on what it observes autonomously.
- The speaker asserts that most people are unprepared for what agentic vision actually means in practice, framing it as a capability shift with broad and underappreciated implications.
- The speaker attributes agentic vision directly to Google, describing it as a new frontier of AI capability and implying it represents an official product or feature release rather than a speculative concept.
Topics
Transcript
[0:00] So, I want to introduce you to a new tech called agentic vision just dropped by Gemini 3 and it's pretty crazy. I mean, it looks at images, it looks at screens, it looks at dashboards. Essentially, it can see. And not only can it see, it can understand the same way a human would. It can decipher, it can plan, it can make decisions based on what it is actually seeing. Which means that this isn't about creating better prompts anymore. If you have an AI out there that is actually observing the real world, digital environments, images, video, and taking action without [0:30] being baby sat, that's pretty crazy. I think we can all agree. And I…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Nick Ponte
This Free AI Agent Can Do $900/Week Tasks While You Sleep (Hermes + Open WebUI)
This video introduces Hermes Agent combined with Open WebUI as a free, locally-run AI agent system that offers persistent memory, autonomous task execution, and over 40 built-in tools. The presenter argues this setup outperforms subscription-based AI tools and can be monetized by offering private AI setup services to businesses with data privacy needs. The video is partly a promotional pitch for the presenter's free AI cash flow masterclass.
Claude Mythos LEAKED - Only 12 Companies Have This Insane AI… Get Ready Now
This video claims Anthropic accidentally leaked information about a powerful new AI model called 'Claude Mythos,' which allegedly found thousands of previously undiscovered cybersecurity vulnerabilities. The video uses this framing to promote a free 'AI Cash Flow Masterclass' teaching viewers how to monetize AI services for businesses. The content blends sensationalized AI claims with a marketing funnel for an online course.
This Free AI Agent Just Got Dangerous (OpenClaw 4.26 Update)
This video covers the OpenClaw 4.26 update, an open-source AI agent with over 346,000 GitHub stars that runs locally and can control browsers, manage files, and execute multi-step workflows. Key new features include durable task flow, multimodal sub-agents, enhanced memory with a 'dreaming' function, and a skill marketplace called ClawHub. The presenter frames OpenClaw as service delivery infrastructure for building recurring-revenue AI businesses.
Google’s New AI Agent Jitro Is Replacing Entire Teams (Here’s How to Profit)
The video introduces Google's leaked internal AI coding agent 'Gitro,' described as an upgrade to Jewels that operates autonomously toward high-level goals rather than waiting for individual prompts. The presenter argues this 'KPI-driven automation' philosophy applies beyond coding and can be used by non-technical people to build income by selling outcome-based services. The video also promotes a free AI cash flow masterclass.
This AI Phone Agent Replaces $3,000/Month Freelancers (Claude + GoHighLevel Setup)
GoHighLevel has added Claude (Sonnet and Haiku) as a selectable AI model inside their voice AI system, enabling agencies to build AI phone agents that reason through conversations, handle objections, and integrate with CRM pipelines. The video argues this creates a recurring revenue opportunity for agencies charging $200–$500/month per client by solving the widespread problem of missed calls and slow lead follow-up for local businesses.