NewsTechnical

AI News: New Gemini Voice + Codex Voice Are Mind Blowing

Paul J Lipsky

Paul J. Lipsky reviews major AI updates from Google and OpenAI, focusing on new voice capabilities in Gemini and ChatGPT for Mac. He demonstrates how context-aware voice features enable screen-aware commands, task delegation, and workflow automation, while also covering OpenAI's significant price reductions for GPT 5.6 models and updates to Gemini Spark and music generation tools.

Summary

Paul J. Lipsky discusses revolutionary voice modalities recently released by Google and OpenAI that he claims have fundamentally changed how he uses his Mac. The core feature is Gemini's new "speak to window" functionality accessed via customizable hotkeys. With the "use reasoning" toggle enabled, this feature becomes context-aware and screen-aware, allowing users to highlight text and request modifications (making it more formal, summarizing, creating emails), process multiple files without opening them, and generate images and infographics based on visible content. The system creates Gemini chats for each query, preserving conversation history. Lipsky provides ten examples demonstrating capabilities from text reformatting to summarizing newspaper articles and drafting responses. However, he notes rough edges where the system conflates dictation with intelligent assistance, sometimes over-delivering when simple dictation is requested. For this reason, he continues using Whisper Flow for pure dictation tasks. ChatGPT's new voice feature represents a different approach: orchestration. Available in the desktop app and controllable remotely via iPhone, it combines natural-sounding voice with ChatGPT's work capabilities. Users can delegate multiple tasks simultaneously, with ChatGPT creating separate chats for each task and managing files, Notion databases, Google Docs, and all available plugins. Notably, users can control their computer remotely via phone, with ChatGPT executing tasks even when the user isn't actively watching. Lipsky organizes his three voice tools by task complexity: Whisper Flow for pure dictation, Gemini speak to window for on-demand in-context assistance, and ChatGPT voice for delegating complex multi-task workflows. The transcript also covers OpenAI's announcement of significant price reductions: GPT 5.6 Luna costs 80% less and GPT 5.6 Terra costs 20% less across API, Codex, and ChatGPT work. OpenAI reportedly used GPT 5.6 itself to achieve these cost reductions. Additional updates include Gemini Spark now supporting Auto Browse in Chrome, ten free Gemini Omni video generations until August 4th, and the release of Lyria 3.5, an improved AI music generator with covers, lip-sync, iOS app, and enhanced musicality controls available through Google AI subscription.

Key Insights

  • Gemini's speak to window with reasoning enabled becomes context-aware and can extract information from files that aren't visually open on screen, as demonstrated when Lipsky had five closed files highlighted and Gemini pulled accurate company retreat information from all of them
  • Lipsky found that Gemini's speak to window conflates dictation with intelligent assistance, frequently over-delivering when he wants simple dictation, leading him to continue using Whisper Flow as his primary dictation tool
  • ChatGPT voice features orchestration capabilities where a single voice conversation delegates tasks to multiple sub-chats that work independently, with the primary chat serving as an orchestrator that can be updated via voice while tasks execute in the background
  • ChatGPT voice can be controlled remotely from an iPhone to manage a Mac computer, allowing users to create notes in databases, conduct research, and delegate work while physically away from their computer, provided the computer is powered on and ChatGPT is open
  • OpenAI used GPT 5.6 itself to reduce the cost of GPT 5.6 models, representing a meta-application where the model optimizes its own pricing efficiency

Topics

Gemini speak to window voice feature with context awarenessChatGPT voice orchestration and remote task delegationScreen-aware AI capabilities and multi-file processingOpenAI model price reductions (80% for Luna, 20% for Terra)Comparison of three voice modalities for different workflowsGemini Spark Auto Browse integrationLyria 3.5 AI music generation updates

Transcript

[0:00] Gemini and ChatGPT have just released new voice capabilities for their Mac apps. And these are some of the coolest and most useful updates Google and OpenAI have ever shipped. I'm not exaggerating when I say that these have completely changed how I use my Mac. I'm going to show you what those new voice modalities are, as well as all the other important AI news from the past week that you need to know about. My name is Paul J. Lipsky and every week I release a video just like this one talking about the AI news that regular people actually care about. So that interests you, make [0:30] sure to subscribe. And now let's get to news.…

Full transcript available for MurmurCast members

Sign Up to Access

More from Paul J Lipsky

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.