This Unknown AI Model is Shockingly Good
A company called RC released an open-source AI model called Trinity Large Thinking that performs comparably to major models like Claude Opus in benchmarks. The speaker struggles to find meaningful ways to test new AI models for everyday use cases since most models already handle typical business and personal tasks adequately.
Summary
The speaker discusses a newly released AI model called Trinity Large Thinking from a previously unknown company called RC. This open-source model, released under Apache 2.0 license by an American company, shows competitive performance against established models including Claude Opus 4.6, Kimi K2.5, GLM5, and Minimax M2.7 in benchmark comparisons. The model demonstrates capabilities in coding tasks, creating games like Snake, and performing what appears to be agentic work through multiple automated steps, though the speaker questions whether the demonstrated coding speed is real-time. The speaker expresses frustration with the challenge of meaningfully evaluating new AI models, noting that most current models already adequately handle typical business and personal use cases that matter to everyday users. They express interest in developing a custom benchmark focused on practical, real-world applications rather than academic metrics, and invite collaboration from their audience to create better testing methods for evaluating large language models as they are released.
Key Insights
- RC's Trinity Large Thinking model performs competitively with established models like Claude Opus despite being from a previously unknown company
- The speaker argues that current AI models already adequately handle most everyday business and personal use cases
- The speaker claims that existing benchmarking methods fail to capture real-world utility for average users
- The speaker believes there is a need for practical benchmarks focused on everyday applications rather than academic metrics
- The speaker suggests that the rapid pace of AI model releases makes it difficult to develop meaningful differentiation tests
Topics
Transcript
We got a new model out of a company that, well, this is my first time actually hearing about them. They're called RC, and they released a model called Trinity Large Thinking. And this is also an open source model released under Apache 2.0, and it's another American company. Taking a look at the benchmarks here, they're comparing it with Opus 4.6, Kimi K2.5, GLM5, and Minimax M2.7. And it seems to be pretty on par with most of those models. They do show it off in use, creating a snake game, doing what appears to be some like agentic work, working through a whole bunch of different steps here. We see it doing some code insanely fast. I don't…
Full transcript available for MurmurCast members
Sign Up to AccessMore from Matt Wolfe
AI News: Dots, GPT-6.1 Sol, Sonnet 5.5, Gemini 4, and everything you need to know
A comprehensive review of major AI announcements from the week, including OpenAI's new Dots agent, GPT-6.1 Soul model, Anthropic's Sonnet 5.5, and Google's Gemini 4 Argon, with analysis of pricing, capabilities, and competitive positioning across different models.
The Hands Down Best Coding Model Right Now
The speaker reviews Opus 5.5, claiming it's currently the best state-of-the-art coding model at a reasonable price point. They demonstrate its capabilities by running a 20-hour test where the AI built a nearly complete recreation of the Mega Bonk game, including accurate character tiers, sound design, and UI elements.
This AI Model Does’t Use Words?!
Typesafe AI released Jev, a new AI model that outputs structured decisions (choices, scores, or booleans) instead of generated text like traditional LLMs. This approach makes Jev dramatically cheaper and faster, with input costs at 4 cents per million tokens and free output tokens.
AI News: Opus 5.5, GPT-6 Sol, Jev, Muse and More!
This week in AI saw major announcements from Meta Connect (Muse agent, new AR/VR glasses), three new language models (Claude Opus 5.5, GPT-6 Soul, Grok 4.7), significant buzz around Typesafe AI's Jev decision model, and new features from YouTube, Microsoft, Google, and Spotify. Claude Opus 5.5 emerged as the new state-of-the-art model while Jev introduced a fundamentally different approach to AI outputs focused on decision-making rather than text generation.
5 Ways To Use ChatGPT's New Features
The transcript demonstrates five practical applications of ChatGPT's new features, including computer use for software automation, advanced image editing, cloud browser integration for account access, voice-powered live agents, and persistent session agents for long-term user interaction tracking.