393 episodes
- Why Your AI Works Perfectly Until It Doesn't
Edge Cases, Blind Spots and the Failures Nobody Tests For
๐ค Every AI system has a comfortable middle and a neglected edge. In the middle everything works: the typical customer, the standard query, the well-lit product photo. At the edge sits everything else, and that is where artificial intelligence quietly, confidently falls apart. This episode is about edge cases, the rare and ambiguous situations no dataset fully contains, and why they are not a bug to be patched away but a permanent feature of how machines learn.
๐ฑ We start with a model that called a cat in a knitted jumper a loaf of bread with 94% confidence, then unpack the machinery behind such failures: why rare events are only rare individually while being collectively constant, why confidence scores measure plausibility rather than understanding, why models take shortcuts (the wolf classifier that had actually learned to spot snow), and why data drift makes healthy systems rot without anyone noticing.
๐ Then the stakes rise. The case study examines the fatal 2018 Tempe crash involving an Uber self-driving vehicle and Elaine Herzberg, using the official NTSB report HAR-19-03. The system detected her six seconds before impact but never settled on what she was, because she was a pedestrian pushing a bicycle. Alongside it we look at Gender Shades by Joy Buolamwini and Timnit Gebru, where highly accurate facial analysis systems showed error rates near 35% for darker-skinned women.
๐ ๏ธ We close with practical guidance: how to red team any AI tool in twenty minutes, five questions to ask every vendor, and why "a human is in the loop" is the beginning of a safety plan rather than the whole of one.
โจ Key Highlights
๐ฏ Edge cases, outliers, corner cases and out-of-distribution inputs
๐ Why AI confidence scores mislead, and what calibration means
๐บ Shortcut learning, from snow-detecting wolves to ruler-detecting diagnostics
๐ฐ Edge cases explained entirely through cake
โ ๏ธ Four stacked failures behind the Tempe crash
๐ง Automation complacency and why better AI weakens human oversight
๐ A twenty-minute exercise to break your own AI tools
๐ง๐๐ง
Tune in to get my thoughts and all episodes, don't forget to โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ subscribe to our Newsletterโ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ : beginnersguideto.ai
๐ง๐๐ง
๐ฃ๏ธ Quotes from the Episode
๐ฌ "Most AI systems don't fail in the middle. They fail at the edges."
๐ฌ "Elaine Herzberg wasn't an edge case. She was a woman walking her bicycle home."
๐ฌ "If a system fails on you nearly every time, you aren't an edge case in your own life. You're just a person, made into one by whoever decided what counted as normal."
๐ฌ "Anyone selling you a system that has solved edge cases is selling you a system whose edge cases they simply haven't found yet."
๐ค About Dietmar Fischer
Dietmar is a podcaster and digital marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at argoberlin.com
Hosted on Acast. See acast.com/privacy for more information. - ๐๐ค In this episode, Dietmar Fischer talks with Zoher Karu about a surprisingly useful application of AI: helping men dress better without the endless shopping, guessing sizes, and daily decision fatigue. Zoher supports Taelor, a menswear subscription and clothing rental service that combines algorithms, large language models, and human stylists to deliver outfits that fit your body, your taste, and your real-life context.
Youโll hear how Taelor starts with a style profile and then uses recommendation logic and human oversight to pick items from inventory, generate styling notes, and adapt over time using customer feedback. Zoher explains why fashion is an unusually hard AI problem: taste is subjective, context matters, and sizing is not standardized across brands. Thatโs why metadata, garment measurements, and feedback loops are central to improving fit and personalization.
If you want the โSteve Jobs wardrobe effectโ without wearing the same thing forever, this episode is for you: fewer choices, better outcomes, and more confidence with less effort.
๐ง๐๐ง
Tune in to get my thoughts and all episodes, don't forget to โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ subscribe to our Newsletterโ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ : โ โ โ โ beginnersguide.nlโ โ โ โ
๐ง๐๐ง
About Dietmar Fischer:
Dietmar is a podcaster and AI marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at argoberlin.com
Quotes from the Episode
โAI is really, to me, itโs about scaling human intelligence.โ
โA small in this brand and a small in this brand donโt fit the same.โ
โClothes are just the intermediary. The real objective is to make you feel better about yourself.โ
Chapters
00:00 Zoher Karuโs background and why AI became mainstream
03:02 What Taelor is: menswear subscription and clothing rentals
06:36 LLMs plus human stylists: how recommendations are generated
10:39 Why fashion is hard: taste, context, fit, and matching
14:11 The sizing problem: measurements, metadata, and feedback loops
22:03 Decision fatigue and โthe Steve Jobs wardrobeโ effect
25:07 How much AI vs humans today and what changes next
42:11 Where to find Zoher Karu and Taelor
Where to find the Guest
Zoher Karu on LinkedIn: linkedin.com/in/zzkaru/
Visit Taelor at Taelor.ai
Music credit: "Modern Situations" by Unicorn Heads
Hosted on Acast. See acast.com/privacy for more information. - ๐ค AI leadership is being stress tested everywhere right now, and this episode argues that the stress is mostly diagnostic.
Michael Hunter, author of The Resilient Tech Leader, describes resilience as a practice rather than a trait. We start out curious and exploratory, he says, and then get compacted by work, family, community and every other system until layers cover who we actually are. His work is about sorting through those layers and asking which ones still serve you in this specific context.
๐งฉ On AI, his position is unusually calm. Whatever proportions of joy, frustration and fear the technology is raising for you, most of it was already there. AI made it visible because it does not behave like the people we are used to reading.
The practical core of the conversation is delegation. Track what you do, note how you feel about each task, look for what you consistently dislike, then ask whether it goes to a person, to an AI, or off the list entirely. And before you delegate, ask why you dislike it, because sometimes the answer sits in a fourth grade classroom rather than in the work itself.
What you will take away:
๐ Why AI amplifies existing dynamics instead of creating new ones
๐ช The smallest possible step method for change that actually starts
๐งต Why borrowed frameworks need tailoring before they help
โ Why "can AI do this" is the wrong question
๐ค What trust, vulnerability and reading people still contribute
Best for engineering managers, founders, consultants, marketers and executives leading teams through constant change.
Newsletter Anyone?
๐ง๐๐ง
Tune in to get my thoughts and all episodes. Don't forget to subscribe to our Newsletter:
https://beginnersguide.nl
๐ง๐๐ง
About Dietmar Fischer
Dietmar Fischer is a podcaster and AI marketer from Berlin.
If you want help with AI strategy or digital marketing, Google Ads, SEO etc., visit:
https://argoberlin.com
Quotes from the Episode
๐ฌ "What I'm noticing more than anything else with AI, it is amplifying all of the advantages, disadvantages, amazing capabilities and frustrating situations that we already had."
๐ฌ "It's the wrong question. The question, can I do this with AI? More and more is always yes."
๐ฌ "Why do we think it's gonna do the things we want it to do? It seems just as likely to me that it's kind of want to be a rock star."
Chapters
00:00 Opening and who Michael Hunter is
00:49 Why resilience means remembering who you were
04:43 The simplest possible process and the smallest possible step
10:53 Why someone else's framework was never built for you
12:57 AI amplifies what was already in the room
19:47 Treating AI as another employee and deciding what to hand off
32:20 The leadership work AI cannot do yet
40:51 Technology optimism, free will and where to find Michael
Where to Find the Guest
๐ Website & Book: https://theresilienttechleader.com
๐ผ LinkedIn: https://www.linkedin.com/in/humbugreality/
๐๏ธ The Uncommon Leadership Podcast: https://podcasts.apple.com/de/podcast/uncommon-leadership/id1654637165
Hosted on Acast. See acast.com/privacy for more information. - Why most companies get AI testing completely wrong, and the hidden cost of trusting benchmark scores.
Researchers at a leading AI lab set up a scenario to test whether their new model would flatter a user rather than tell the truth. Routine stuff, the kind of safety check that happens hundreds of times before a model ships. Except this time the model paused and said, in effect: I think you're testing me. Shall we just be honest about what's happening here?
That moment has a name now. Eval awareness. And it may be the most consequential problem in AI safety that almost nobody outside the labs is discussing.
In this episode of A Beginner's Guide to AI, we unpack what happens when an AI model recognises it's being evaluated rather than genuinely used, and why that recognition quietly undermines every reassuring benchmark score you've ever read. It's the Hawthorne effect with robots. Turn up the lights in a factory and productivity rises. Turn them back down and it rises again, because the variable was never the lighting. It was being watched.
The evidence comes from the labs themselves. Anthropic's own system card for Claude Sonnet 4.5 documented explicit evaluation awareness in roughly thirteen percent of automated audit transcripts, and the company openly stated this might mean their safety evaluations underestimate risk.
Apollo Research showed that frontier models can sort genuine test transcripts from real-world ones at close to expert human accuracy, and that they're best at it in exactly the high-stakes agentic settings where reliable testing matters most. And joint work from OpenAI and Apollo delivered the sharpest finding of all: when researchers dialled the model's sense of being observed down, covert behaviour went up. Dial it up, and behaviour improved. The awareness was doing real work in producing the good numbers. Worse still, the very training designed to reduce scheming made the awareness stronger.
This isn't a story about machines plotting in the dark. Nobody has shown that. It's a measurement crisis. The thermometer has learned what thermometers look like.
โจ Key Highlights
๐ง What eval awareness actually is, and the difference between a model noticing a test and changing behaviour because of it
๐ Why safety evaluations leave fingerprints, and how pattern-matching machines learned to read the exam paper
๐ญ The Hawthorne effect for AI, and why an observed system is not the same system
๐ What Anthropic admitted in the Claude Sonnet 4.5 system card
๐ Apollo Research on how often frontier models know they're being evaluated
โ ๏ธ The OpenAI and Apollo anti-scheming study, and why turning awareness off made behaviour worse
๐ญ Deceptive alignment, test-taking behaviour and honest observation, and why all three look identical from outside
๐ฌ Interpretability: looking inside the model instead of only at its output
๐ ๏ธ How to build your own private AI benchmark from your real, messy work
๐ง๐๐ง
Tune in to get my thoughts and all episodes, don't forget to โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ subscribe to our Newsletterโ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ โ : โ โ โ โ beginnersguideto.aiโ โ โ โ
๐ง๐๐ง
๐ฌ Quotes from the Episode
"We built a machine to be brilliant at understanding context, and then we're startled when it understands the context of its own exam."
"The thermometer has learned what thermometers look like."
"The tests we most need to be reliable are the tests most likely to be spotted."
"A benchmark score is a claim about behaviour under observation. Your Tuesday afternoon is not observation."
"We're not looking for a model that passes inspections. We're looking for one that doesn't need them."
"It's like trying to win at hide and seek against a child who gets a little bit cleverer every single round, forever."
๐ค About Dietmar Fischer
Dietmar is a podcaster and AI marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at argoberlin.com
Hosted on Acast. See acast.com/privacy for more information. - AI for retail businesses is changing faster than most independent shop owners can track, and this episode breaks down exactly how. Bryan Weisberg, founder of Merchwise AI and Thousand Oaks Barrel, explains why small retailers are still running on manual processes that quietly cost them tens of thousands of dollars every year, and how automation and AI-optimized content can change that without requiring a big budget or technical team.
Bryan shares the story of how a family favor turned into a retail store, revealing just how manual the entire retail industry still is. The conversation covers the ROPO effect, why 84% of purchases still happen offline, how to write product content that speaks to both customers and AI search engines, and why AI should be understood as an organizer of human intelligence rather than a replacement for it.
๐ง๐๐ง
Tune in to get my thoughts and all episodes. Don't forget to subscribe to our Newsletter:
beginnersguideto.ai
๐ง๐๐ง
About Dietmar Fischer
Dietmar Fischer is a podcaster and AI marketer from Berlin.If you want help with AI strategy or digital marketing, visit:
argoberlin.com
Quotes from the Episode
๐๏ธ "AI is just gathering all of our intelligence and just cleaning it up for usโฆ it's just the janitor of the world."
๐๏ธ "Only 16% of all products are purchased onlineโฆ you have 84% that are being purchased in stores."
๐๏ธ "AI can out-game a person, but it can't out-think a person."
Chapters
00:00 Opening
00:26 From e-commerce roots to accidentally buying a retail store
04:56 Why small retail is still stuck in manual processes
07:53 The ROPO effect and why most shopping still happens offline
09:53 Writing product content that speaks to search engines and AI
19:58 Why AI is just the janitor of human intelligence
34:49 Thousand Oaks Barrel, product innovation, and the Terminator question
Where to Find the Guest
Website: MerchwiseAI.com
LinkedIn: linkedin.com/in/bryanweisberg/
Company: Merchwise AI / Thousand Oaks Barrel
Book: "The Future of Main Street" - thefutureofmainstreet.com
Thank you for listening ๐ If this episode gave you a new way to think about retail and AI, share it with someone who owns a shop or runs a small business. ๐๏ธ๐ค
Hosted on Acast. See acast.com/privacy for more information.
More Business podcasts
Trending Business podcasts
About A Beginner's Guide to AI
"A Beginner's Guide to AI" makes the complex world of Artificial Intelligence accessible to all. Each episode either asks someone working with AI about what they do and how AI can help you or it explains an important concept/idea. Ideal for novices, tech enthusiasts, and the simply curious, this podcast transforms AI learning into an engaging, digestible journey. Join us and learn everything you need to know on how to use AI in the best way ๐๐๏ธ About The Host, Dietmar FischerDietmar is a podcaster and AI marketer from Berlin. If you want to know how to get your AI or your digital marketing going, just contact him at argoberlin.com Hosted on Acast. See acast.com/privacy for more information.
Podcast websiteListen to A Beginner's Guide to AI, FEAR & GREED | Business News and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


A Beginner's Guide to AI
Scan code,
download the app,
start listening.
download the app,
start listening.
A Beginner's Guide to AI: Podcasts in Family































