PodcastsHear the voice. See the shape of the thought.

#fable-5#anthropic#llm-benchmarks

We Tested Anthropic's Fable 5 for a Week

Dan Shipper, CEO of Every, spent a week with Fable 5 — Anthropic's Mythos-class frontier model — before its public launch and walked away genuinely changed. Every's senior engineer benchmark put Fable at 91/100, against 63 for Opus 4.8 and 62 for GPT-5.5 — a jump Dan describes as "warp drive" capability for sustained autonomous work. The model is slow, expensive, and token-hungry, but for anyone orchestrating big, multi-hour agentic tasks, there's nothing close to it right now. ## [00:00] One prompt built an infinite 3D library Dan opens with a live demo: a fully browsable 3D version of Jorge Luis Borges's "The Library of Babel" — hexagonal galleries, accurate mathematics from the story, working bookmarks — all generated by a single prompt. He gave Fable a one-line instruction to read the story, plan, and execute a browser-playable 3D game end-to-end. The model ran autonomously for three to four hours, self-checked its work, and shipped. > *"I made this entire thing in a single prompt with Fable 5, the new model from Anthropic."* ## [01:22] Our day-zero Fable 5 review Dan introduces himself and Every's approach: they test models hands-on for real production work — programming, writing, design, business decisions — and report back on what actually works. Fable generated unusual levels of pre-release hype; Anthropic had initially said it was too dangerous to release. After a week of internal access, Every's take is that the model is genuinely different, and Dan's goal here is to cut through the excitement and show the realistic picture. > *"Because we've been using this model for about a week now, we get to pull back the curtain a little bit and show you what it's like to have lived with this model."* ## [02:25] What a Mythos-class model is Mythos is Anthropic's new top-tier model family, sitting above Haiku, Sonnet, and Opus in their lineup. Architecturally it's not novel — same transformer family, just bigger. Anthropic added strict safety guardrails (no cyber, no biological use cases) to make it releasable. Pricing is steep: $10/M input tokens, $50/M output — roughly 2× Opus. Dan's verdict from a week of use: genuinely the most powerful coding model he's ever touched, by a wide margin. > *"It is just genuinely the most powerful coding model I've ever used by far."* ## [03:28] The 91/100 engineering benchmark Every runs a proprietary senior engineer benchmark: the model is handed a real "vibe-coded slop" production codebase and asked to rewrite it from first principles as a senior engineer would. Prior to Fable, the top score was Opus 4.8 at 63/100, with GPT-5.5 right behind at 62. Fable scored 91 — matching a human senior engineer in a single prompt. Dan had expected saturation of this benchmark in about six months; it happened in two weeks. > *"Fable scored a 91 on this benchmark. 91 out of 100. That's the same score as a human engineer with just one prompt. That's crazy."* ## [04:12] Why it feels like a warp drive Fable's core strength is sustained autonomous execution over multi-hour tasks. You give it a destination, leave it running, and come back to something finished. Unlike earlier Claude models that eagerly said yes to everything ("purple accents, purple accents"), Fable deliberates, pushes back when something can't be done well, and follows through on complex, loosely specified prompts. Dan's analogy: a warp drive — not instant, but it compresses what used to take months into hours. > *"You can specify a destination for a big trip, and it just compresses what normally would have been like years or months into like hours or days."* ## [06:10] Where the model falls short The warp drive metaphor cuts both ways: it's useless for getting around town. Tight back-and-forth collaboration, quick questions, rapid iteration — Fable is a poor fit for all of these. It's slow, expensive, and burns tokens aggressively. A non-obvious workaround: drop the reasoning level to medium or low for simpler questions; that's how Anthropic's own people use it internally. Without a big, meaty problem to throw at it, the model is overkill. > *"If you're using it for true collaboration or quick questions or things that need tight back and forth, I don't think it's that good for that."* ## [07:04] Building a Heidegger lecture site Dan describes asking Fable to grab philosopher Hubert Dreyfus's 2007 lectures on Heidegger — without even providing a URL — and turn them into a consumable mini-site. Fable found the lectures, wrote per-lecture summaries, built a synchronized player that highlights the transcript as audio plays, added chapter navigation, drop caps, and typographic choices that Dan characterizes as actual taste, not the default template output. One prompt, no scaffolding. > *"That's what I mean when I talk about this model having really exceptional taste and attention to detail."* ## [09:05] Finding a growth bet in customer data Every has ~10,000 paid and ~100,000 free subscribers and a backlog of survey data the team had been analyzing with AI for weeks without a sharp conclusion. Dan fed it all to Fable. In one pass, the model came back with: "You have a conversion merchandising problem. Your free-to-paid conversion ratio is lower than it should be." Then a falsifiable bet: ship pricing transparency and a trial offer, and it'll go up. That synthesis — reading survey responses, site analytics, and product state together — hadn't emerged from weeks of team analysis. > *"That is something that I would expect a really, really good growth person to do with a lot of time and thought and research."* ## [10:35] Clearing a real GitHub backlog Every's agent-native markdown editor Proof accumulates GitHub issues automatically as agents file bugs during use. Dan pointed Fable at two weeks of open issues and told it to close irrelevant ones and write Rust fixes for the rest. It swept through the backlog and produced patches the team actually merged. Other models can do this, but they require hand-holding — one issue at a time, constant check-ins. Fable just batched it. > *"And it just went boom boom boom boom boom boom. And actually wrote fixes that we merged."* ## [11:17] Who should actually use this model Dan is direct: Fable is not for everyone right now. Using Every's "eight levels of AI adoption" framework, it pays off at levels 7–8, where users are already orchestrating multiple agents and have large problems queued up — typically technical builders. For knowledge workers not yet running agent workflows, it'll feel like overkill; for casual vibe coders, the token costs are real friction. About half of Every's own early-adopter team saw immediate payoff; the other half is still growing into that workflow level. > *"Using it is a skill. You need to be exposed to problems and working at a level of expertise where the problems come up in order for it to be useful."* ## [13:31] Where other models still win Writing is the clearest gap: Fable's prose is dense, literary, and block-heavy — good for thinking through structural writing problems, not for copywriting or everyday sentence-level work. For Claude users, Opus 4.8 is still better for writing. For GPT users, 5.5 is a better daily driver. Dan himself keeps GPT-5.5 as his Codex driver for the quick back-and-forth that fills most of his day; Fable gets reserved for big production pushes. > *"For my day-to-day, it's a bit overkill even for me."* ## [14:26] What this means after automation Dan points to his essay "After Automation" as the frame: automation doesn't shrink human work, it creates more of it — a paradox. Fable follows the same pattern: it raises the floor for non-experts (a vibe coder can now one-shot a video game) and raises the ceiling for experts (an expert can build a AAA game solo). The displacement is real and he says it's normal to feel unsettled by it — but the capability curve means even people who can't afford Fable today will have access within six to twelve months. > *"This model increases the floor of capability for non-experts, but it also raises the ceiling for experts."* ## [16:02] The final verdict Dan closes with a straightforward recommendation: read the full Every vibe check for detailed benchmark breakdowns across coding, writing, and knowledge work, watch "After Automation" for the bigger-picture framing — and then go find the first big problem you've been avoiding and point the warp drive at it. > *"If you're psyched about this, the thing I recommend most is go use your new warp drive. And let me know what you make."* ## Entities - **Dan Shipper** (Person): Co-founder and CEO of Every; sole presenter in this episode; spent a week testing Fable 5 pre-launch. - **Every** (Organization): AI-native subscription media company focused on testing frontier models for real work use cases; ~10,000 paid subscribers. - **Fable 5** (Software): Anthropic's Mythos-class frontier model; scored 91/100 on Every's senior engineer benchmark at launch. - **Anthropic** (Organization): AI safety company; maker of the Claude / Opus / Fable model family. - **Mythos** (Concept): Anthropic's top-tier model family tier, above Haiku, Sonnet, and Opus; characterized by extended reasoning and high token cost. - **Senior engineer benchmark** (Concept): Every's proprietary evaluation — model rewrites a production codebase from first principles; scored out of 100; Fable hit 91, Opus 4.8 hit 63. - **Opus 4.8** (Software): Previous Anthropic flagship; scored 63/100 on Every's benchmark; still preferred for everyday writing tasks. - **GPT-5.5** (Software): OpenAI's comparable frontier model; scored 62/100 on the benchmark; Dan's personal daily driver for quick back-and-forth work. - **Hubert Dreyfus** (Person): American philosopher; author of "What Computers Can't Do" (1972); subject of the Heidegger lecture site demo. - **Proof** (Software): Every's agent-native markdown editor; used in the GitHub backlog-clearing demo. - **After Automation** (Concept): Dan Shipper's essay arguing automation creates more human work rather than eliminating it; referenced as the interpretive frame for Fable's broader significance. - **Eight levels of AI adoption** (Concept): Every's framework for classifying AI workflow integration depth; levels 7–8 are where Fable delivers the most value.

The SaaS Apocalypse Is a Goldmine With Figma's Matt Colyer

33:53

#saas#ai-agents#developer-tools

The SaaS Apocalypse Is a Goldmine With Figma's Matt Colyer

Figma developer PM Matt Colyer has been building his own AI agents for two years and is buying more software subscriptions than ever — not fewer. He and Every CEO Dan Shipper work through why the "SaaS apocalypse" narrative gets the economics backward, how AI needs to escape the tyranny of the text box to unlock genuinely creative design work, and why the coming year's challenge isn't generation but review: humans are now the bottleneck in a world where agents can ship faster than anyone can evaluate what they made. ## [00:00] AI will create a billion developers This exchange, taken from later in the interview, opens the episode: Matt argues that the number of developers worldwide — roughly 25–40 million a decade ago — is heading toward a billion. That demographic explosion, not AI replacing software, is what makes the SaaS market a "gold mine." Figma and most established SaaS businesses are, in his view, excited rather than threatened. > *"If you're in that space, like, it means it's a gold mine, right?"* ## [01:03] Introduction Dan Shipper frames the conversation: he recently bought Figma stock after noticing the "SaaS apocalypse" discourse, and he wants to know how a company that pre-dates AI is navigating a world where agents can now operate inside your product. Matt, as the director managing Figma's developer products, is the right person to ask. > *"There are all these people who are like, 'Oh, I don't have to use Figma anymore.' You guys just launched an agent in your product. You also have Figma MCP."* ## [02:15] Why the SaaSpocalypse narrative has it backwards Matt's counter-argument runs on two tracks. First, the democratization of software creation massively expands the addressable market — more software being built means more demand for the tools, infrastructure, and services that support it. Second, vibe-coding your own app sounds liberating until you're dealing with SMTP upgrades at midnight. He built his own email agent two years ago and watched it get rickety; these days he pays someone else to run agents for him rather than maintain the plumbing himself. > *"I'm buying more software these days than I ever did before, because I'm like, 'You know what? That tool seems cool. I'm just going to pay somebody else to run my agent for me.'"* ## [05:27] Matt's email agent origin story The origin was unglamorous: three kids in three schools, relentless PTO emails, and the humiliation of missing spirit day. Matt wired up a Python script to grab his inbox and paste it to an LLM — the whole thing was rickety and sometimes the replies didn't work, but the core loop worked. He then added a memory system and a daily summary pushed to him proactively, which he flags as the real unlock: instead of having to open a tool and ask, it just showed up. Dan mirrors this with his own Codex-based inbox workflow, now four weeks into inbox zero. The two also land on voice as an underrated interface — Matt uses Loom recordings because it feels less weird than talking to a blank screen. > *"The unlock for me was like instead of having to go to a tool and ask for the thing, it was just like it would show up."* ## [13:21] Divergent vs. convergent design thinking Chat-based AI is inherently linear — you iterate on one design thread. Matt's argument is that great design has a diamond shape: first you diverge (generate many directions), then you converge (pick the best). Figma's on-canvas agent is a first attempt to break out of the text-box constraint. On the canvas, an agent can spawn a grid of frames — grayscale, sepia, with different type — and then a separate convergent agent can cluster them and recommend which direction to pursue. Command-line agents can't do this kind of spatial, parallel exploration; that's what the canvas unlocks. > *"Text boxes are super limiting — it's very much like a linear 'well this and then that.' If we get to the canvas, the agents allow you to do divergent thinking."* ## [17:39] Figma's MCP server MCP gives third-party agents (Cursor, Windsurf, Claude Code) a standard interface into Figma. Two flows: code-to-design — fire up a dev server, ask the agent to screenshot a live page and pull it into a Figma canvas — and design-to-code via "Get Design Context," which wraps component properties and design library guidelines into an agent prompt that then creates a branch, writes the code, and posts a screenshot to the PR. Both flows remove the manual copy-paste drudgery that used to live between the design file and the codebase. > *"You pull up your codebase, fire up the MCP server, and ask it, 'Hey, can you go to this page and copy it into Figma canvas?' And it will actually do it. That's a little bit mind-blowing."* ## [19:45] Why design agents need personalization Generic agents produce generic output. For Figma, the difference between an okay agent and one people actually love is whether it understands the design system — the components, the spacing rules, the naming conventions. Without that personalization layer, generated designs aren't usable. Matt draws a parallel to the memory systems in chat agents: in Figma's case, the design library is the memory. He also hints at proactive agent work Figma is cooking internally, framing the core problem as maintaining design values at a pace agents can generate. > *"The thing that really differentiates an okay agent from one that people really love is the personalization aspect. For Figma's version of that, it's the design system."* ## [22:09] Every problem is a context problem Matt describes a Figma product operations team that realized every recurring PM task — onboarding docs, project tracking, team introductions — was a context problem in disguise. They built "PMOS": a local SQLite org chart wired to Asana, Slack, and GitHub, then layered Claude Code skills on top. When a new team member joins, the system walks the org chart, reads the last 30 days of Slack channels, checks the Asana board, and produces an uncannily good onboarding file. Dan points out that Claude Code's power comes from the same insight: instead of an always-on cloud agent you have to manually wire to everything, it's an agent that already has access to everything on the user's machine. > *"One of the unlocks to me about AI is like you kind of realize every problem becomes a context problem. The work becomes about framing the problem with the right set of information."* ## [25:12] Apple and Google as the reigning kings of context Matt has been waiting for Apple Intelligence to deliver on its WWDC promise — phones hold all the personal data; an always-on, actually-smart Siri should be the obvious product. It hasn't arrived. He's watching Google's rumored "Spark" agent (always-on, connected to all Google content) with similar anticipation. Dan's take: Apple wins regardless because everyone runs AI on Mac hardware, giving them time to catch up. Matt adds that Apple's privacy-first positioning is a genuine strategic asset, not just PR. > *"Even being late to the game, they are still the king of context. And I think that's what's been interesting to watch about Google I/O this year — seemingly Google has also kind of woken up to that."* ## [28:18] Why review is the new bottleneck Generation is no longer the hard part. Agents are cheap, capable, and available; the problem is that humans are now inundated with net-new content they need to evaluate and approve. Matt frames "review" as the coming year's core design challenge: how do you scale a human value system — what good looks like, what fits your brand — at the pace agents can ship? The format is still unsettled: video walkthroughs, screenshots, a trusted review agent. He closes with a thought on careers: fundamentals still matter (you need to know what long division is even if you use a calculator), and the people who will thrive are the curious ones who ask how something is put together rather than just accepting the output. > *"We have agents that are capable of producing all this stuff, they're available enough, they're cheap enough. We're just being inundated with new content. The bottleneck is now: how do we scale our value system to evaluate it?"* ## Entities - **Matt Colyer** (Person): Director of Product Management for Developers at Figma; has been building personal AI agents for two years; longtime developer tools practitioner. - **Dan Shipper** (Person): Co-founder and CEO of Every; host of the "AI & I" podcast; active AI agent practitioner (inbox zero via Codex). - **Figma** (Organization): Design and prototyping platform; launched an on-canvas agent and an MCP server; central example in the SaaS-in-the-AI-era discussion. - **SaaSpocalypse / SaaS Apocalypse** (Concept): The narrative that AI will make SaaS software obsolete; both guests argue the opposite — AI expands the developer population and demand for SaaS. - **Diamond-shaped design thinking** (Concept): Divergent phase (generate many options) followed by convergent phase (select the best); Colyer argues current chat-based AI only supports linear/convergent work. - **MCP (Model Context Protocol)** (Concept): Standard interface for third-party agents to connect to tools like Figma; enables code-to-design and design-to-code workflows. - **Figma MCP Server** (Software): Figma's implementation of MCP; supports live page screenshot-to-canvas import and "Get Design Context" design-to-code export. - **Claude Code** (Software): Anthropic's coding agent; referenced as an example of an agent with full local file system context; used by Dan Shipper for inbox management. - **Every** (Organization): AI-focused media and software company; Dan Shipper is co-founder/CEO; runs the "AI & I" podcast series. - **Proactive agents** (Concept): Agents that push summaries or actions to users without being asked; Matt identifies the proactive daily email summary as the unlock that made his agent genuinely useful. - **Review bottleneck** (Concept): The emerging constraint in AI-assisted work where generation is fast but human evaluation/approval capacity is the limiting factor.

10:30

#claude#opus-4-8#llm-benchmarks

Why Opus 4.8 Pulled Me Back to Claude

Dan Shipper, CEO of Every, delivers a day-zero vibe check on Opus 4.8, arguing Anthropic could have called it Opus 5. The model jumps 30 points past Opus 4.7 on Every's Senior Engineer benchmark, edges out GPT-5.5, tops their internal writing tests at 79.6 vs. 73, and is the first model to produce a genuinely good one-shot slide deck. Two catches temper the enthusiasm: performance degrades sharply below "extra high" reasoning, and the Claude desktop app remains cluttered compared to Codex. ## [00:00] What is Every Every is a 30-person applied AI lab for the future of work—part media outlet, part product studio. Dan opens by explaining the subscription (writing, courses, AI-built tools all in one place at every.to) before rolling into the Opus 4.8 assessment. The plug is brief and context-setting: the team has had beta access for a week, and the rest of the video is what they found. > *"Every is the only subscription you need to stay at the edge of AI."* ## [01:07] Anthropic Is Back: The Headline Case for Opus 4.8 Dan had largely abandoned Claude after Opus 4.7—slow, hard to love, and outpaced by Codex and GPT-5.5 in day-to-day use. Even the most loyal Claude users at Every had started routing work elsewhere. Opus 4.8 breaks that pattern: it scores 63 on Every's Senior Engineer benchmark (30 points above Opus 4.7, one point above GPT-5.5), tops their writing tests, and produced the first one-shot slide deck Dan has called genuinely good. Kieran Klaassen, Every's GM, called it "the most human model he's worked with." The one persistent friction is the Claude desktop app itself. Codex is fast, focused, and ships a clean harness; the Claude app still feels like a product built by three separate teams—chat tab, code tab, co-work tab, each with its own feel. Dan is now splitting time between both apps, which he was not doing before. > *"But honestly, they could have called it Opus 5 cuz this is a really great model."* ## [05:02] Reach Test: Paradigm Shift Ratings from the Every Team Every's reach test asks one question: do you actually open this model when work gets hard? Dan rates Opus 4.8 gold/green—paradigm-shift quality, docked one notch because the Claude app harness is only "okayish to pretty good." Kieran, who runs 50 agents a day, gives a straight gold paradigm-shift, one of the rarest grades the team has assigned. Katie Parrot, a senior staff writer and historical Claude fan, lands at green, splitting her work between Opus 4.8 and Codex. > *"It's very rare to give a paradigm shift grade to a model. So I would pay attention to this."* ## [06:32] Benchmarks: Coding and Writing Numbers On coding, Opus 4.8 hits 63 on the Senior Engineer benchmark—the test feeds the model a vibe-coded codebase and asks it to rewrite from first principles, then scores against two human senior engineers who completed the same rewrite (typically scoring in the 80s–90s). GPT-5.5 sits at 62. On Kieran's LFGbench (real-world tasks: SaaS build, e-commerce site, 3D game landscape), the model writes readable code that bridges technical competence and creativity—the "cozy island" 3D scene is notably richer and more vibrant than GPT-5.5's output. On writing, Opus 4.8 scores 79.6 out of 100 on Every's internal benchmark (intro writing, promo emails, mid-piece paragraphs); GPT-5.5 scores 73. The gap is mainly in AI tells: at high and extra-high reasoning settings, Opus 4.8 produces prose that sounds less like a model. It matches a writer's voice from a single paragraph of context better than any other model Dan has tested. > *"Opus 4.8 scores a 79.6 out of 100 on the writing benchmark. GPT 5.5 is 73."* ## [08:57] Emotional Intelligence, Knowledge Work, and the Verdict Dan uses the model for interpersonal and management work—talking through decisions, pressure-testing his own framing. Opus 4.8's thinking traces show it genuinely cycling through permutations before responding, which makes it feel less like a sycophant and more like a useful counterpart. On knowledge work, it's versatile: code and writing coexist cleanly in a single thread, and the slide deck result is the first one-shot deck Dan would actually send to someone. The verdict: if you're a Claude fan, this model delivers. If Codex converted you, add Opus 4.8 as a parallel tool for writing and knowledge work—it's worth the context switch. The harness gap is real, but the model itself is a banger. > *"If you've been converted to Codex, I highly recommend you at least add it as part of your arsenal."* ## Entities - **Dan Shipper** (Person): Co-founder and CEO of Every; presenter and primary evaluator of Opus 4.8. - **Kieran Klaassen** (Person): GM of Kora at Every; gave Opus 4.8 a straight gold paradigm-shift rating on the reach test. - **Katie Parrot** (Person): Senior staff writer at Every; rated Opus 4.8 green, split between it and Codex. - **Every** (Organization): Applied AI lab and media subscription company focused on AI for the future of work. - **Anthropic** (Organization): Developer of Claude and Opus 4.8. - **Opus 4.8** (Software): Anthropic's latest Claude model; subject of the vibe check. - **GPT-5.5** (Software): OpenAI model used as the primary performance comparison across all benchmarks. - **Codex** (Software): OpenAI coding agent; praised for its clean desktop harness and used as the daily-driver counterpoint to Claude. - **Senior Engineer Benchmark** (Concept): Every's proprietary coding benchmark—rewrites a vibe-coded codebase from first principles and scores against human engineers. - **LFGbench** (Concept): Kieran Klaassen's real-world coding benchmark covering SaaS, e-commerce, and 3D scene generation tasks.

Wir haben alles mit KI automatisiert und unsere Mitarbeiterzahl verdreifacht

41:13