chapters transcript notes
click any line to jump to that moment in the video
0:08 Hello everyone and welcome back to Agentic Thinking with Mike and Matthias. We're back again for another great topic. We're going to talk about some news items today. The main topic for day is the Vibe coding trap. 0:21 Did that you were walking into a trap when you were building Vibe coding things? I think this is a relevant article. This is about AI and how it's potentially bankrupting early startups. 0:32 this is an interesting article. We're going to unpack this today. Matthias, how are you doing? I'm doing great and yeah, excited to see what we come up with here. 0:47 We may have different opinions. Let's see what happens. That makes just for good tv, you have different opinions about things. let's step into this one. 0:58 without a whole bunch of ado here, let's just go into the main topic here, which is this article which I will put. It's a Medium article and Matthias, I don't get to see all of the article because it is a subscriber only article here. 1:11 if you want to go check this out, this is on Medium is the article that we're reviewing here, you should go check it out. The article is in the description, it's also in the chat window as here. And I'm going to lead you, Matthias, to give us 1:23 a little bit of a runup here on what's in the article and why is this topic coming up ? What's happening with startups with AI? 1:33 What's the trend here that's happening? Go ahead and take it away. Yeah, I found that article and I thought that would be a great one for us to go through because it's opinionated. 1:43 It resonates with me saying businesses building on top of Vibe coded software and wipe coded products 1:58 apparently increasingly find themselves in quite a pickle. they definitely manage very to, , to produce hype and to, 2:10 to deliver and ship something that may look appealing. But when it comes to building a business on top of 2:23 such a software, there may be some traps, some bigger ones, some, some smaller 2:33 ones. And the article has a few very nice examples, , around database corruption, missing database indexes, about 2:47 massive waste of tokens for AI applications, , when, when they don't really engineer what they're sending back to the LLMs. 2:58 And plus obviously a number of security issues, , around 3:11 protecting, , anything you, you put put out as, as a public website or . obviously some of the Stories are pretty 3:24 drastic and, probably represent extreme poor and bad examples. But nonetheless, I think there's a lot 3:37 here to unpack in terms of our own experiences as . What have we seen, how do we feel about it? And also, is it to think of a dichotomy of vibe 3:52 coding on one end of the spectrum versus, I don't know, engineering software engineering on the other end of the spectrum? I'll leave it here for . 4:03 Yeah. this is interesting to me. I believe. Microsoft did a very cool demo, I think in the March timeframe. They had some users on a stage and they're 4:13 talking about , , how do we apply AI, where do we apply it in our report writing or data engineering process? What is this? What does AI look in the hands 4:23 of a new engineer versus a seasoned engineer, which one would win? And they pitted these AI pieces a novice engineer with AI in their hands and said, hey, go build this 4:34 report, go build this data. It's up to you to use the AI with you to go build something the business may want. And then they gave the same task to an engineer who, who 4:46 knows the program, knows how to build things, and said, go build this equivalent information. you got these two worlds of novice with AI and expert without. 4:56 And they did this demo of between. we're going to give you two report pages. Which one was built by the AI, which one was built by the person, which one was built by the novice or the expert. 5:07 And it was really interesting to watch. It was somewhat hard to determine in certain scenarios, other places it was very clear which one was which. But at the end of the day, they were trying to bring 5:18 the parallel that bringing AI to a novice user helps you build incredible things much faster. But you lose some of that expertise, that knowledge around the experts in this space, which is throttling capacity management 5:33 using only the amount of information you need, things that you kind of learn the hard way of building them on your own. And when you give AI to an expert who, who's a 5:43 really good developer, you get a compounding effect of really good output from them. They're able to work faster. All these edge cases. And I think partly to this article, which 5:54 is, , it's alluding to, I believe, partly here around the single, it's the millionaire solopreneur, ? 6:05 I'm a single shop, I'm building my app, I'm developing it, and then I just launched it on Production day. And it goes until the SQL database falls over. It goes until the API starts throttling. 6:16 And all these users who are trying to use your app, it's falling over. had you been an engineer designing this, you probably would have gotten to the same end result for an 6:26 app, but your infrastructure on the backend side would use some caching. There'd be different API patterns, there might be different architectural decisions that was being made to scale this to high volume. 6:37 Let me just pause there. Is this where you feel this article is going, Matthias, or is there a different vibe that this article is pushing in? 6:51 one thought that I just had as you were talking, and I think that's where the article is going, I'm not 7:01 decided yet. And maybe we can discuss that whether there is a really strict distinction between 7:12 a seasoned engineer architecting something and then building a system with the help of AI on the one hand, as opposed to, you call it a novice user or someone who doesn't 7:25 approach that from an engineering standpoint and who gives all the powers to the AI. Vibe codes the whole thing, ? AI makes the decisions. 7:38 The article obviously says if it's the latter, you're bound to get into some trouble. maybe minor, maybe major, may be bad that ultimately 7:51 it bankrupts your business because, , you may have launched but you haven't been able to then scale the thing or secure it or make it as cost efficient as it should 8:05 be. . what I'm wondering is if we're saying there is this inherent danger and threat and challenge from wipe coding, is that something 8:21 which is just a matter of time in the sense that, you know, as new generations of models come out and as the tools we're using to do the vibe coding get better is, , 8:35 are we just waiting for some iteration of AI where those problems no longer exist? Or do we genuinely or always will have a 8:48 distinction where AI and LLM driven development can only go far? And we will always need, , for something that's really thorough 8:59 and scalable and secure and solid, will we always need a trained, experienced engineer there? . I think ultimately that's the big question we're faced with today. 9:13 And I'm not sure I'm keen to hear what you think, but I'm not sure anyone has a definitive answer yet. Let me add a parallel of something I've heard messaging messaged over 9:25 the last couple months. Let Me see, I may take a tangent here, Matias. This may not be a tangent. I don't know. There are some really strong experts in the Power 9:36 BI community, namely SQL BI and Tabular Editor, which I think are very great teams of people, extremely smart individuals, no discrete to any other acumen 9:50 in the Power BI data modeling in space. I feel I've watched a shift occur where and I don't 10:00 think they were saying, to be very clear, these teams of people were not saying no, don't use AI. I think the initial rounds of AI interaction from Tabular Editor and 10:11 SQL BI were , look, hold off. The models are not quite good enough to handle what we're trying to do. There still needs to be a human in the loop. Don't just brain dump and let the AI try to build decks 10:24 and models and everything just willy nilly. ? I think that part was clear. But I feel recently, in the last two to three weeks 10:34 I've seen an immense amount of recommunication around those teams shifting heavily into using AI. Tabular Editor is coming out with a 10:44 CLI tool that's helping it become much more capable for building tymdal and semantic models using agents. Much more structured information. the Tabular Editor team has built a harness around agents that 10:58 are making it very friendly to model development. I think this is something they had to do because Power BI had come out with the MCP modeling server. Why would we need Tabular Editor if you have an MCP modeling 11:10 server and I can talk to an agent directly? We've done this on the show, We've done a number of demos here of talking to the agent model and say, build me a diagram, change these measures, group these things together. 11:19 Why do I want to go talk to another third party tool where I can directly talk to an agent, have it done faster, not having to click any buttons. I think there's been a shift here where these teams have had 11:29 to greatly adapt this one. I'm seeing advertising from Marco Russo and SQL BI saying, hey, we're going to teach you how to use DAX with agents. How does agents write dax? 11:39 What do we need to communicate to an agent that it's good at writing dax? These were not messages I was seeing three months ago, four months ago. And all what I'm trying to 11:51 say here and maybe bringing it back to your point here around where are things going? The models are getting much better. The harnesses that people are developing are getting much more capable. 12:01 The large language models can Hallucinate can do weird things. But what you want is you want a very accurate system around a harness, wrangling that large language model to allow it to reason, 12:14 to allow it to build things that are useful. But it needs to be able to test. It needs to have specifications. These have linting built into it. Did it even write the syntax that goes into the DAX 12:26 statement or not? . these are things, I think, that are trying to be organized in the community . And we're seeing a shift where more and more experts are adopting 12:37 AI and building tools around it, which is making it more useful for everyone. back to this point here. This is where we are today. 12:47 I think these models are going to continue to get better. The harnesses will continue to get more capable. People are going to continue to build better tools that are going 12:57 to enable better usage of this moving forward. And while I think we're going to have some users here. . I'm going to rapidly build an app. It may fall over, it may break on the back end. 13:11 . Shame on me. I'm a new user to Vibe coding this whole application thing. But it doesn't take long for you to realize, oh, it wasn't that the AI couldn't build, was more that your 13:24 prompts and what you described to the AI was insufficient. And I think that's much more of the scenario here, which is 13:35 we just have people not being able to prompt correctly into the AI to get out what they want. And I know this because, Matthias, you're putting me to shame. 13:48 I'm doing a certain amount of programming and Vibe coding and building apps, and I probably wouldn't call it Vibe code anymore. I would really want to call it, , agentically coding, because I know what I want to build. 13:58 But, Matthias, you're building things above my pay grade, way better than what I'm doing. Much more elaborate and complex. But you're using agents to do a lot of the development. 14:08 But that's because your system that you've been working on and developing is just done. And it's incredible, , what it can do. And I think the more the general population and public 14:20 starts learning what this looks in people, building programs and applications that help this to become more real, it's just going to shift my work away from clicking buttons and moving things around to describing 14:33 things to an agent and having it pair program, or pair build with me alongside here. let me just pause. I'm gonna just leave that thought there for a moment. What do you think about that? Is that, is that alignment with where you're thinking things are going 14:43 to. I think one thing we can take for granted is today, with the harnesses and models we have available to us, you 14:56 can build very, very, very substantial, complex tools and software applications. . If. 15:10 . Big. If what to ask for. If you do 15:21 a fair amount of the architecture upfront and then feed that into the prompt. As you just said, prompting is what it's all about. 15:35 Yes, but are we ever going to get to a point where you're going to get similar outputs from an AI without having to do that 15:48 upfront work? That's really what I wonder about. . Because that would mean that we no longer have a wipe coding trap to go back to the article. 15:59 . And I don't think we will ever get there. If you want to believe all the big players and the anthropics 16:11 and openais in this world who want to sell you their tokens for lots of money, obviously they make you believe that you just 16:22 do voice input to a dot or a bot or whatever you want to call it, and then they build the next million dollar 16:32 business for you. Yes. Without you holding any hands. . Yeah. I am very, very, very skeptical about that. 16:42 But on the other hand, I'm extremely excited about the capabilities we have with the current generation of models, 16:53 allowing you to build, architect and build very substantial systems and applications 17:04 for very, very cheap and very, very quickly, , . Yeah. Do you just pushing it back to you. 17:16 Do you think I'm wrong with my skepticism? Do you think this is just a matter of evolution? 17:26 I, I think we're in the middle of the evolution process of this. . I, I think if we look at, if I look at where things have come from January till , 17:37 in January through maybe March, I was doing a lot of experimentation. I was building very small, rough versions of applications. I feel I'm spending a lot more time polishing and 17:49 getting something that was way closer to a real production item that I would want to put my name on and put out the door as branded items. 18:00 I only see this progression moving further towards. It's going to get more capable, we're going to get better models, we're going to get more use out of these models for us 18:11 as we build things. I feel the progression is moving towards. In the same way I made the analogy on the Explicit Measures podcast, in the same way that Power BI reports were democratized 18:22 for the Business user, the BI team, we're democratizing application and app development by using agents to do the same thing. I don't have to go into a Power BI report and ask 18:33 or figure out how to write D3. The visuals all run on D3. I just take for granted that there's a framework around it. And when I click the button to make a bar chart, boom, 18:43 bar chart shows up, I drag my fields, boom. It all maps together. Someone spent a lot of work around this to make it an easy experience to use. They democratize all the code that was required to write that. 18:57 In the same way, I see agents being that Power BI tool that democratizes the app. 19:07 Because some users are able to build apps and scale them in systems, ? You may not be able to and that I think 19:17 that's more of a limitation on your knowledge around what does an at scale app look ? How do you communicate to the edit? for example, when you're doing a planning mode around 19:28 what your app may be or what design you may be building, you say, look, I need this app to be able to be scalable, ? some of the prompt that I should be presenting to it is this app needs to be scalable. 19:37 It needs to have these types of design features on the back end. You can describe more of those features, the AI will return. Hey Matthias, for building your app, did you need a 19:49 redis cache? If you don't understand what a redis cache is and why that helps your SQL Server not fall over or your API not throttle because you're hitting the same request over and over again. 20:02 If you don't understand that, then how can you tell it, yes, I need it or no I don't, ? You. what is really is interesting here is you can say things 20:14 you, you have to learn how to use language of , I'm going to have again, it's, it's good, clear requirements, ? When I'm building the app, I got to be able to tell my, even my developers today, ? 20:24 If I told you, Matias, I'm going to Build app for 10 people, you would have in your mind a certain architecture. And then you would ask me questions around , , how often 20:34 is the data going to be accessed? How much data are we accessing? Is this just a little bit or a lot of bit, ? That will change part of how you design the solution. But you already know this because of your historical experience. 20:44 You've gone into these problems and ran into a wall and how to solve it. The AI knows how to solve these problems. I'm just not telling it the requirements. 20:55 And I think that's where things for me, I'm , I'm looking at this and saying, I think the AIs are really good and they can outbuild what I am aware of or what I'm knowledgeable about. And I have to direct the 21:08 agents. And one area that I'm going to say is very hard to do, and even I struggle with this sometimes My team asks me questions from the AI because they'll talk 21:20 to it and they'll have an answer and they'll respond back. It gives you two options that are very close in capability. 21:30 . How do I have a preference to push the AI one direction versus another? . What, what knowledge or background information do I have that helps me 21:42 design the solution for what I want? And sometimes the AI will come back to you with two choices that are probably both good in their own . But maybe a downstream design solution feature that you want to build 21:55 may conflict with these two items or this decision point. . And. And that's where I think the expertise of the expert starts coming into play. You're able to zoom out a 22:05 little bit more, look a bit. A bit of a big picture, figure out what's going on, and then zoom into the very specific details of what the agent's asking for and then catering your prompt better to what they need. 22:17 I'm hearing you say yes, it, it feels maybe this is something you've encountered with your agents where it's giving you two good options, but you got to pick one. Oh, yeah, yeah. That's definitely something which resonates. 22:29 Generally you get more than two. And I generally, my experience, you generally have a recommendation as . 22:40 When there are choices to make. Oftentimes the agent highlights what they would recommend, but I recognize what 22:51 you say. Oftentimes I would then push back and say, can you present the options in a different language? Can you dumb it down? 23:02 Or can you explain what the impact will be? things that, and that normally helps. Or another way to go would be for me to say, 23:16 I don't care. I've put you in charge of delivering this goal. You make the call. As long as you achieve the strategic goals I set out with you in the beginning, you are free to 23:27 do whatever you want. , that would be another approach. I think ultimately it's about being selective. what component do you care deeply about that you 23:41 get very hands on and that you get very involved with the architecture and what component doesn't need that where you where you 23:53 can afford to hand it off to an agent in terms of making decisions. . That's interesting. I that idea of there are definitely parts 24:03 of the application or whatever you're building where you're going to be very stringent on. it has to be this way. This is the vision I have. Yeah. , , identity authentication, payments, . 24:15 Yes. that hard currency stuff. Yes. But do not mess it up. UX or Docs or. 24:31 Yeah, anything. Anything that's more cosmetic rather than ultimately it's, , again, coming back to the article which obviously 24:41 says if you're just wipe coding, you may be able to start a business, but , don't be surprised if it goes bankrupt. . At the end of the day it's, it's, it's a 24:53 risk assessment. . And you want to think about what's the risk of the AI getting this particular piece wrong. 25:05 If, if that's a risk you're not willing to accept, then you better get more hands on here and you better get more involved. . I that part. Yes. I think it's a good, I think 25:16 it's a good litmus test of , , if the AI broke this part, would it be a problem? again, to this article going back maybe here a little bit to this one and talking more about it, which is, , 25:28 where does the artist, where does the, the author here start picking on, hey, where should we be putting our pressure? ? Some of our pressure should be on the database design, 25:38 the API design, telling the agent to build tests around high volume capacity. . this is, these are part of the techniques I think that we would apply here and from those techniques that 25:51 would then elevate more of our design. . We can still rely on the agent to build it, but telling it better requirements will change the design that it is 26:03 using caching. It is not letting the SQL database fall over. And maybe to some degree here, unless there's, most of this is, 26:13 I feel most of this is fixable. if you, if you did go down the route and you're gonna, you're gonna burn some users and , the people are gonna adopt the app and the app starts crashing, , fine. 26:26 There's very little that you can do that is unrecoverable. At least quickly building through . Hey, agent. My SQL database is falling over for these reasons. 26:37 What's going on? That's a really good. that the problem, you go back to the agent and say, build the monitoring, build the error messaging that we need 26:48 build. And you can rapidly iterate through these problems. And while it was quick to build the app, you also have another agent. I think people forget this. There's an assistant that's there to help you debug them too, and 27:01 figure out better architectures. And as fast as it was to create the original app, you probably could recode that with another agent to fix it. 27:11 Most of it is recoverable, except for one thing, and that's very. I want to say, security here. No, no. . Business model. Oh. 27:21 If you made assumptions and launched a business, , with. With a. with a pricing model or . Or with certain contract types. 27:31 Yeah. And then it turns out that is never going to be profitable. Yeah. You're in the pickle. And the article has some examples for that as . . That's a great. 27:41 That's something you can't really recover from if you're charging too little for the app you're using in. The app is very expensive. It'll never work. 27:51 Which, if you don't mind, leads us quite nicely to open AI, which we covered last week because they've just had a similar realization. 28:03 they've. They've had a. A $200 max subscription. Yeah. For a long time. Which they put on pause at some point. 28:15 And last week they launched a new 500 plan, , bigger than any of their. Of their competitors. 28:25 Yeah. And they also brought back the 200 plan, but with substantially reduced quotas in. In within the tour. 28:37 Sorry, I think I'm about to. Good. Sorry, I was. Didn't want to sneeze into the microphone. they got a lot of heat for that from people who. 28:50 Who were on the $200 plan suddenly being told, , you can stay on it, but we're reducing what you're getting out of it by 50% or . 29:01 It's definitely not going down very . But that is an example where something that happened. And obviously we've seen that with respect to the cost of using 29:11 AI getting increasing continuously. That's something we talked about a few times. June 26th was a big deal when Copilot completely changed their pricing 29:25 model. Companies that can get away with that to some degree, but you are small business. May not be. 29:37 This, I think, is way more impactful because that you have this, I feel they're slowly cooking the frog here in this example, 29:49 ? You're, you're getting value at 40 bucks a month. great. We're, we're starting to throttle you. We're starting to get some limitations. You're getting too many tokens used, ? 29:59 we're going to move you up to $100 a month plan. , great. everyone starts doing $100 a month plans. . And then we have this transition to. in the business world, , we're going to transition you 30:11 away from monthly user plans and we're going to transition you into pay as you go plan. there's this whole new realm of pay as you go and tokens become more important. And we're getting even more premier 30:23 models. And we've gone from 200amonth all the way up to 500amonth for these plans. And , if you're a team of three people, , 200 to 500, yes, it's a hit on 30:34 your business line. But if you're providing value and adding value from that, , not that major of a hit. If you're a Microsoft, if you're a SAP, if you're a big 30:45 organization and you're trying to, and you're trying to figure out how to go from $200 per user to $500 per user to get them to do what they're doing and you're cheapening out some of 30:55 these lower end plans and they're getting less tokens, they can't get things done. You're starting to feel a pinch. This is boiling the frog. We're starting to bring the price level 31:05 up for these higher amounts. And the more you increase the price on these plans to get more tokens out of them, I think the harder people are going to search for optimization, 31:19 doing things efficiently, picking the model for the solution. this is what brings to mind the hydrofusion that, that GitHub is coming out with. hydrofusion is going to be Microsoft 31:31 solution to this, which is picking the model mid prompt. How do you pass context to agents or different models when they need to be in Hydrofusion? these are, these are areas that 31:42 I think are going to be more important moving forward. 31:52 Yeah, sorry, just got distracted. No worries. Do you have the 500amonth open AI plan? Are you needing that one? 32:03 I don't know. I don't need that. I use a combination of plans from different providers and 32:17 I use Some very dynamic routing, let's put it that way, when it comes to choosing the provider, the model, 32:29 the effort level. Yeah, that definitely took a lot of engineering and experimentation as because I don't wasting money. 32:42 I'd, I'd rather keep the bill low, . Absolutely no. I never even got to the 200 plan either. 32:53 I don't need it. They've also had, interestingly, , ever since the 33:05 big announcements from open air last week, interestingly, they've had quite some trouble regarding performance in particular. 33:16 it's definitely something which I've been following quite closely because I use codecs and open AI. Sure. A lot of my day work, if 33:27 they go down, you can't use them. You're seeing that. I've definitely noticed massive performance degradations with respect to 33:39 speed and with respect to, , token flow. Not necessarily around quality. Although. Although, , if you go online 33:52 you'll find quite a few people who claim that anything below Astra is no longer as smart as it used to be. 34:02 I can't really judge that. I. It's not necessarily my own experience, but there have been quite a few outages where mid session models suddenly report that they're at capacity 34:17 and you need to retry and things that. It's not looking great, let's put it that way. Yeah, 34:29 it feels when these large models are announced and everyone gets really excited about these new things, there's this level of usage that happens, spike of usage that happens here that really pushes people over 34:39 the hill. we hope open I, open AI can figure themselves out here. one interesting article today. Go check out the article if you can on medium. It's really good. It's written by Abdul and I recommend go reading 34:53 that one. The Vibe coding trap why AI generated MVPs are quietly bankrupting early stage startups. there's some definitely considerations there as you educate yourself on how you build apps and build things there. 35:05 That also being said, thank you Matias for a great conversation around OpenAI and what's happening with their world and how we're seeing and using AI in our workflows day to day. 35:15 That being said, thank you all much and we'll end here with a little treat for everyone as we end our episode today. Thank you all and we'll see you next time. 35:25 Thank you and thanks everyone. 35:47 Jack from Finance builds his first report at night. Experts might but his team sees the light. 36:01 Five minutes to wow when the room starts to leave. Power be I rising 36:13 from the ground unseen. One shore before the tower bow. 36:29 Jack from fin is pouring joy into. 36:50 Talk about crease can take a year to land. Bottom up. Belief it's in someone's head. 37:03 A colleague shares a fire, the spark goes wild. Adoption isn't force, 37:15 it's arising time. 37:41 Power . They said Surprise would choose another name 37:58 but the people on the desk still played the game because the two felt close 38:10 because the wind fell near bottom up that's how the truth appears. 38:21 Don't wait for the mandate from a distant floor. Start with the person who needed one door. Give them five minutes and a reason to believe the rest of 38:33 the company love us to breathe. There's a night festival the base under my chest. Think of Jack and how the quiet ones progress. 39:27 Sam. 40:02 Agency thinking, Agency thinking.