Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Devoxx · @DevoxxForever
Words
7,971
Runtime
44:53
Speaking pace
178wpm
Reading time
33min
178 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[Music] systems are nominal initialize Genesis sequence welcome everybody thank you for attending my talk especially while there are so many of you really great um we're going to talk about the intersection of tdd and gen today and we're going to see if this is a perfect pairing so I would like to start with an introduction why am I here talking about this subject so last year I had a talk uh called the battle of the AI coding assistance I gave it at multiple conferences multiple
89 words, the words spoken in the first 30 seconds at 178 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1 |
| Average words per sentence | 7971.0 |
| Longest sentence | 7,971 words |
| Questions asked | 0 |
| Sentences containing a number | 1 |
Most used terms
Filler phrases
167 in total: um 82 · uh 53 · like 26 · actually 4 · kind of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
[Music] systems are nominal initialize Genesis sequence welcome everybody thank you for attending my talk especially while there are so many of you really great um we're going to talk about the intersection of tdd and gen today and we're going to see if this is a perfect pairing so I would like to start with an introduction why am I here talking about this subject so last year I had a talk uh called the battle of the AI coding assistance I gave it at multiple conferences multiple meetups also in this one last year and uh in this talk I showed off features of different AI coding assistant think um get up co-pilot jet brains AI assistant but also good old jet GPT and uh one of the features that I showed off was the fact that they are pretty good at generating tests how does it work you as a human you vary implementation you ask the AI to come up with tests and in most cases it comes up with a happy case in most cases it comes up with a few Ed cases and if you're still not happy you ask for more so this worked really well after one of those conferences somebody came up to me and asked why don't you do it the other way around and I said the other way around what do you mean by the other way around and she said now what if you as a human righty test and then ask the AI for an implementation I said wow that's an interesting idea I don't know if it works but I'm going to investigate and I'm going to get back to you this investigation led to this talk so the second thing I would like to do is in the section of the title so let's have the title again the title is TD and geni are perfect pairing there are three it Concepts in here and the first two are plain obvious the first one of course is tddd test driven design and the second one also really obvious gen but there is also a third it Concept in there who dares to guess wow tough crowd okay cool the third one is pair programming so that was hidden in the subtitle so what we're going to do we're going to take those I3 three it Concepts and we're going to see if we can combine them into something new but before I'm going to do that I first going to take a few minutes to talk about those three concepts or we're all on the same page so I would like to start with geni and um jpt is the most famous one it is introduced two years ago it feels like so much longer but it's already been only been two years with us um and since then a lot of changed has changed because every big um every Big Industry Giant in the it has their own jetbot nowadays and they all have something in common they all run on a technique called large language models and those models are pretty good at generating text and also pretty good at generating code and that's the feature that we're going to use today then the second it concept is test driven development so people hold entire conferences about this subject but I'm going to boil it down for this presentation to write tests in advance and this is better than doing it afterwards for multiple reasons but two reasons I really like and the first one is the fact that if you first write your implementation and then your test tests your test will be biased by your implementation and we do not want biased tests secondly um I really believe the following the more time you think about something the better the end result will be and when you first write your test you think a lot about your implementation and therefore you will write a better implementation so those are the things I really like about tdd and it works really well with pair programming so let's talk about pair programming so pair programming is really simple you write software in pairs so you sit together behind one computer and you write software normally you do this um we have two rules so we have the driver that's the one sitting behind the keyboard writing the software and we have the Observer or Navigator and that's the person who's giving live feedback to your code and that's the person you're discussing with while you're writing the code also a thing that you see a lot while doing pair programming is the fact that one person writes the test you switch chairs and then the other person writes the implementation so let's see if we can combine those three concepts into a New Concept can we create a perfect tripling so what we're going to do we're going to pair program with Gen the human will write a test and the AI will write try to write the implementation I call this test driven generation so actually test have a double purpose in this idea and um I talked a lot already about the first purpose tests are the input for this idea so you as a human right to test we give them to Ai and we ask for an implementation but there is a second purpose and the second purpose is probably its most important purpose because the second purpose of test is checking if the output of the large language model the implementation is a correct implementation and this is really important because by now every body has been using llms for two years and you all know that those llms sometimes tend to make things up and that's a problem especially when we asking the llm to come up with a proper implementation um and to illustrate this point I'm going to tell you a story about Air Canada so there was a story in the news in the beginning of this year about Air Canada which is an airliner and they have a website and on their website was a chatbot and uh a customer was talking to the chatbot no sorry I'm telling it wrong let's start over so they have a chatbot and Air Canada decided to improve the chatbot by adding a large language model to the chatbot and this in itself is a proper idea because then your chatbot will act not act but will the sentences from the chatbot the responses will be more natural more humanlike um so the following happened a customer was talking to this chatbot and as you probably already imagined the chatbot started to make things up and they were talking about about refund policies and the chatboard started to make up refund policies the customer was foresighted made screenshots of the discussion and then a few weeks nothing happened but after a few weeks the customer called the help desk with actual humans behind it and he said I would like to refund my ticket and the people behind the help Des said sorry that's not possible he said yes that's possible look I have screenshots the the chatbot told me it's possible I can refund my ticket and the people said no the chatbot made a mistake but the curer disagreed and went to court and the judge totally agreed with the customer the judge said I don't care uh if a human or a chatbot put something on your website if it's on your website it is true therefore you have to refund the ticket the next day the chat bot is offline so I think this story really illustrates the problem we nowadays have with large language models so they're pretty good at the first 95% but the last 5% is the important part we need output of large Lang language models that we can trust for the full 100% and lucky for us we developers we have a way to check the output of the large language model because we have our unit test so we run the unit test against the output of the large language model so my name is Balon I'm the CTO at a company called syn ICT and together we going to see if tdd and gen is a perfect pair so I created little schema to show off our way of working and I'm going to start with the happy flow so I'm going to start on the left and then we go straight to the right afterwards I will discuss the loops in the model so on the left we start with test tests and those are input for the Gen we asked the Gen to come up with an implementation that passes the test um remember this is Happy flow so we get a perfect implementation and that means all tests will pass that means we follow the yes arrow and we're done if you trust your test you have production ready code unfortunately in the real world things go wrong and that's why we have loops in the model so we have a small Loop underneath gen uh I call this the code loop and a big loop uh on top I call this the test Loop so the small one the code the code loop uh is necessary because sometimes gen comes up with an explanation instead of an implementation that doesn't work sometimes gen comes up with partial code or non-c compiling code not good enough you cannot run tests again that against that so then the top Loop the test Loop uh sometimes not all test pass if that's the case we extract the error feed it back to the Gen and ask for an implementation that will also pass the test that failed just before and this is how we get a feedback loop so um demo time so what I'm going to show you I want to show off the process so we're going to start really really simple uh so I created a test called the OD even test let me show it to you and as you can see it is really really simple it's a junit 5 test uh it expects an object called OD even to be there a class or even I should say and this class has one method called is even if you provide it with a even number it should return true and if you provide it with an odd number it should return false so really really simple don't worry I will step up the game during the presentation there will be more interesting examples this is just to show of the procedure and um I have a a choice for the audience um in my left hand I have the blue pill which is jet GPT and in the right hand I have a red pill which is clawed so who prefers jet GPT raise your hand Okay who wants me to do the demo with Claud oh wow Claud easily wins okay cool we're going to do CLA what I'm going to do I'm going to copy the code I'm going to go to Claud I expected Claud so thank you and I also have that gbt so I am prepared um so we're going to ask please Implement a class that passes this test and and I'm going to paste the code I'm going to press enter and now we just wait for a second we're on conference Wi-Fi so it's always a bit of a of a thingy but hey look at that we get an implementation back and as you can see it looks like a proper implementation but luckily we have test to check if it's really a proper implementation so I'm going to copy this going to go back to my ID I'm going to go to the test subject notice that I have a dumy implementation here with an incorrect incorrect implementation I'm going to select everything I'm going to paste it I'm going to go back to the test and I'm going to run the test and we hope for a green check mark and look at that green check mark okay this works but okay this was a really really simple example I I understand why you're not standing on your chairs and applauding we're going to step up the game so the next thing we're going to do is prime number generation and again I have a test for that that prime number this is the test again it's pretty simple simple test um we expect a class called prime number generator to be there um it has one method called generate it takes an integer as uh input this integer denotes the upper bound of the prime numbers to generate and if you ask for 50 we expect this answer so doable right again I'm going to copy everything I'm going to go to CLA normally it's pretty smart I just can paste the code and it remembers from the previous uh question that I want an implementation and look at that there's an implementation we're not going to spend too much time on reading the implementation because we have tests to see if there are correct so I do exactly the same again Dy implementation select everything paste the code back to the test run the test and let's hope for a green check mark again a green check Mark okay this also works I'm really happy the principle works thank you for coming to my presentation enjoy the remainder of the conference no no no no just kidding thank you thank you for the Applause um I'm really happy that the principle works but I have two major issues with the fact with the current solution and the first one is really simple there are too many manual steps and what do we we programmers do when we have too many manal steps we automated thank you it's like you read the slides so uh let's automate this um and just to uh explain why we should automate let's rehash the steps I just did I copied code from my IDE I went to a browser I typed in a prompt I paste the code I pressed enter wait for the answer extracted the code went back to my IDE went to the test subject pasted back to the test run the test press to creen Triangle wait for the outcome check the outcome only talking about this makes me tired so we really really have to automate this and um so that's what I tried to do and that sounds simple right you saw me just doing it I was in my IDE went to my browser typed a little bit wait copied a bit how difficult can this be um so yeah it was a little bit harder than I expected I expected me to take three to five evenings to build a tool that would automate this for me H yeah in the end it took me 10 to 15 evenings and a few weekends but I learned something I learned that the IDE does a lot of work for you and of course we already know that the IDE does a lot of work for us but in this project it became really apparent and I will Circle back to this in a few minutes in the end I had a tool called the test driven generator and that's the tool we're going to play with in a few minutes but before we going to play with it I first want to take you through all the little steps that it takes so probably you recognize this as part of the schema that I've shown you before this is the left part and um the first thing that you have to do you have to call an API of a gen um I started building this tool in the beginning of this year and I started out by using jgpt because that's the one I was most familiar with at the time uh later in the process I also switched over to other ones uh to see which ones work better this is not hard but this work and you have to call an API you need a key all kind of work you have to do but hey it's just work one evening then you have to interpret reply and that sounds really simple because you just saw me doing it in a browser right it was just a popup and I could just copy the code if you use different llms interpreting the reply isn't as easy as it seems because all the llms use a different way to denote what is code and what is text so this is more difficult than you think then we have to create a file um and this file needs to have the proper file name because as you all know the Java if you use Java the file name should be equal to the public class name so we need a feature to extract the public class name from the code and then we have to create a file then we have to create this file in the right location because if your code contains a package keyword the part behind the package has to be equal to the folder structure where your file uh resides so we have to be able to extract the package part and we have to able to create directories then we go to this part then you have to compile both files I I didn't see that one coming I'm using intellig as my IDE I never think about compiling files anymore who thinks about compiling files the IDE does it for you and it just works always at least 99.9% of the time then I have to load both classes uh the test file that I use as input but also the implementation from the llm didn't think about it this is typically something that ID does for you then I had to add the test Library so normally when I start a project I add junits in the p and I forget about it but now you have to load junit in the tool otherwise it cannot run test didn't think about it then you have to run the test from code now I always just press the green triangle but apparently you can also run test from code and it's not difficult but it's again work you have to do once we all the test pass I copy the implementation back to the project where the test file was and so you can directly use it so then the the big loop the test Loop um if not all test pass you have to extract the errors um not difficult but hey you have to dive into memory structures extract which test failed and then feed them back to the Gen um while I was doing the demo you saw me the second demo with a prime number generator I just pasted the code and Claud was smart enough to remember that it had to write an implementation that's because if you do things like this in a browser it has context it remember remers what you talked about before this is not the case when you use an API I didn't know this so you have to build a context mechanism yourself and lastly I started out this project with a few system out print lines like every project um but every time when you talk with an llm as you all know llms generate a lot of lot of text um I needed a proper locking solution to keep track of what was happening inside the tool so I just told you I started out by using jet GPT but I also heard more and more about local large language models so I decided let's try those and they were easier and faster than I expected I expect them to be really hard to set them up like 20 manual steps and a lot of aror prone steps turns out you just download a tool and then you run a model they were also much faster than I expected um I expected answers to take up to minutes at that time I had a an M1 MacBook and answers to about tens of seconds this is an M3 and it even generates answers under 10 seconds so it's pretty pretty fast there are lots and lots of local coding LMS llms out there um but in the end I settled on llama 3.1 nowadays there's also 3.2 but 3.1 still performs pretty well for this presentation so that's the one I'm going to use for the next demos when I'm using a local llm so finally demo time so what we're going to do first I have to do a get reset I also forget so that's why that's what thing is there so I'm getting rid of all the things we just did with cloth I'm getting rid of all the ansers of cloth everything is clean again and um what we're going to do we're going to do the same thing prime number generation but now without all the manual steps so in its current form the tool is just a jarf fall in a different directory so you can just call it and what it needs it needs a path to the test file so you see the test file on top I'm going to copy the path and I'm going to paste it over here and what you're seeing is we're in the first test Loop within a test Loop you can have five code loops and of both you can have five so it does 25 iterations at Max this is configurable uh and then it tries to find an answer now this one is pretty simple as you can see it found one test that's the one on top of the screen and this test succeeded we can take a look look at the implementation but we don't have to because if you trust your test you trust the output of this tool and this is really great remember me doing this by hand a lot of steps now I can just write my test and start a tool grab a cup of coffee and come back and hopefully I have an implementation like in this case so the tool works too and that's really great let's take some time to see um what the implications are of using a tool like this so how would you add features so in the past I would write more algorithms right I would grab my implementation add some lines of codes and I would have new features but the implementation is now the field of the AI I'm not touching the the implementation anymore I'm only allowed to add to write tests and that's how you add features you just write more tests and that's what we're going to do so remember the test case with prime numbers we're going to add a test case for negative prime numbers so let's go over there so remember this is where you could specify the upper bound of the prime numbers to generate if you put in a negative number over here um you're asking for negative prime numbers as far as I know they do not exist so I want this code to throw an exception so how do you do that you just add another 10 so like this I think this is really funny it's incorrect but the fact that I have an under 50 and the AI thinks I want an over 50 I think that's pretty cool unfortunately it's incorrect so I have to fix it negative input then the AI thinks I need an object called prime number generator totally correct so pretty happy with myi and this is incorrect but hey still 50% um what we need is something called an assert rows so this one so look at that AI exactly created what I want I'm going to explain it so with an assert throws you first specify the exception that you expect so I'm going to expect an illegal argument exception secondly you provide a Lambda that should um throw this exception so this is how an assert throws works so I'm going to run over here exactly the same command I didn't change anything on the terminal um the difference is in the fact that it still points to the file on top of the screen and that file contains now two test cases you see I found two test and both tests succeed now let's have a look at the implementation to see what changed and this is the new part so if you want to add features to your code you just write a new test and that feels really really natural so what about existing code so all the demos I just gave we created something new but I assume most of us are programmers so raise your hand if you're a programmer yeah very good most of us and uh you all wrote hundreds maybe thousands tens of thousand unit tests in your life so you could also use this tool um and feed it existing unit test and and see what happens maybe it comes up with something better than you created do you dare to try it I did um actually I use the tool to improve the tool and that's something I see more often with AI talks um in the tool is something called the code container and it has a few features the first feature it has to hold the code that the llm produces secondly it has to be cable to able to extract the public class name and transform it in an file name so we can write the correct file and the second feature is it should be able to extract the package part so we can use it to come up with the right folder structure so I have a test called the code container test and let me show you what it does so this test checks the file name extraction part so I have a few classes simple classes over here and notice that I put you have a normal class but you have public classes random spaces in spots you see like this and this um um yeah let me tell you okay so I've I started out with a really crappy string manipulation solution um if you have some experience with programming uh most people would solve extracting something from this uh code by a regular expression but hey I have a confession to make I'm pretty bad at regular Expressions so I started out with this crappy string manipulation solution um and that worked until I started adding test cases with those spaces in between then everything broke every time so let's see if the AI can do better so this is the input then this is what I expect and then lastly there's a message I think adding a message is always a good practice when writing test not only for yourself but for the person after you who will use the test but in this case it also really helps the AI because it has more information it knows what it did right wrong and you give it extra information to come up with a better implementation and we do something similar here for package name extraction so notice that we have all kind of packages over here notice that we have single level packages multiple level packages random spaces again in spots and also a class without package because package of course is optional in Java and again you see here the answers at the end and some messages to help the AI or the next programmer who will touch your tests um what we're going to do we're going to copy the name of the path to this file I'm going to go over here I'm going to get rid of the last input I think it's this shortcut yes it is and I'm just going to paste uh the file so let's see what happens so the tool is now talking to llama 3.1 running locally on my machine uh trying to come up with an implementation that passes this test and um yeah that probably will take some time now this is interesting what what happens is we go to the second test Loop that means it didn't come up with a proper code there was no compiling code in which I could run tests against and it doesn't look well for oh hey look at that it came up with an implementation um it found two test but the implementation didn't pass any test this is a nice moment to show of the the loging so over here you can see I'm using AMA on my local machine it's us using llama 3.1 and I'm using a time out of 30 seconds and as you can see it failed no solution found so what do you do when your tool doesn't work what is the first thing we can do we retry I'm not sure if anybody said that but that's what I do so that's the first thing we do but um because we all also know about llms that they sometimes hallucinate sometimes they're stuck in a corner and they unable to get out yeah look at that in six seconds one shot solution um and this isn't really a simple uh it took me some time to WR this class uh and this is not this is not a trivial um problem anymore so let's see if we can have a look at the implementation so we have a code container uh which uses regular expressions like I predicted as you can see it came up with a proper regular expression mechanism so I'm pretty happy with my tool it saves me a lot of work and it comes up with better code than I did maybe that says more about myself than about you but still still still pretty happy um at the beginning of the talk I told you I'm pretty happy with this tool but there are two major issues and the first one is we needed to automate so that's why I created the tool but secondly um what about libraries and Frameworks so now I have a tool that's pretty capable of coming up with plain Java code which is nice but that is not really what we do at work right most people in this room will use libraries or Frameworks and um the tool in its current form can cannot handle that so just doesn't work with the Jaron and that's also diff difficult problem to solve because the input for the jar file is a test file and from the test file it's pretty difficult to extract which libraries or Frameworks are used so I started thinking how can I solve this problem and the answer is pretty obvious you have to do something with the build system so in my case it's Maven and Maven knows about all your dependencies so therefore it also knows about the libraries and Frameworks that you're using so I decided to create a maven plugin uh I use Maven because I'm more familiar with Maven than I am with Gradle but the principle holds and should also work for Gradle um now the plugin knows about your dependencies because it's a MAV plugin I expected it to load all dependencies because I never wrote a the M plugin before um but this hypothesis turns out to be wrong it does not load all your dependencies but the fact that you are a MAA plugin makes it really easy to load all dependencies so I wrote a custom class loader and that loads all the dependencies that it can find in your pom.xml so um let's demo this as well I've created a simple test called the endpoint test and actually it is is a spring boot hello world test that's what it does so uh for people who don't not know spring boot test they look like this most cases they have a random Port which you can use over here sorry about that so we have a test rest template because we want to do a rest call to our spring boot application and we do a get on the root and we expect it to return hello world now let's uh um run the tool I told you it is a mava plugin therefore we have to take a look at the pom file which is over here and uh there's a lot of stuff in there but the most important stuff is over here there's a new plugin called the test driven Maven test driven generation Maven plugin and then we have configuration and it's most important input is the fact that you can provide a PA to a test file the same as the jar file and this file is the file I've just shown you so what we're going to do we're going to run Maven test driven [Music] generator just generate like this and oh wow I think I made a typo let me see what I did wrong I think it's test driven generation sorry about that generation like this yeah as you can see you recognize this uh those two lines it just reuses a jarf fall I've used in the previous previous demos of course and um currently it's talking to my local llm llama 3.1 and it doesn't go really well so I just told you the first thing that you do when something doesn't work is retry so that's what I going to do first but if that also doesn't work um we're going to start using a cloud llm because I'm really Happ happy with the local llms but the cloud ones are still so much better than the local ones uh so that's something you can do and um there's a small bug in my tool sometimes it doesn't directly Release Control you see there is no terminal I cannot type so unfortunately it shows up during the demo so sorry for that but we can just retry it over here so Maven test driven generation generate so we retry um and we just wait I'll give it a few seconds if it's still not okay now I did something that's already much better you see it found one test and one succeeded so give it a little bit more time but what we can do yeah we can I'm already going to prepare the P XML so we can run to a cloud uh model and it is really simple you see some configuration over here so what I'm saying is use the entropic family you can then also specify the model but there all kinds of settings you can set in your pom file um I'm going to assume this is not going to work so I'm going to kill it and I'm just going to rerun it over here so the thing that I changed is this line and if I now go to the lock file generator dolog you see it goes to entropic and it use claw 3.5 soet and look at that build success in 8 seconds so let's take a look at what it created let me first make some space I have to do that over here then we go to here it created an endpoint application uh which is a spring boot application it is a rest controller and now uh for people who know spring boot this is how you start it uh a get mapping on root and it returns hell world this is exactly who how I would have written it myself so pretty well done uh uh May plug in but okay I agree this was again a small example let's step it up a notch so what we're going to do I also created an endpoint h test so let's take a look at that test endpoint h test so um pretty much the same the idea here is we provide an endpoint with a birth date and we expect it to come up with the age of a person and let me walk you through the examples so if you're born on the first year of this year you're zero if you're born on the first day of 2020 you're four if you're born on the first day of 2000 you're 24 but if you're born on the last day of 2000 you still have to book come 24 and you're 23 so this is the idea um and notice that the age should be part of the URL so slh SL and then the birth date now um I go to the P file I change this path course I should use the h test I go to my terminal and I rerun and I have to warn you I would say in 15% of the cases the local llm can oh I'm talking about it's talking to a cloud llm now it doesn't matter that save some time 50% of the cases my local llm can solve this but this is on the edge what a local llm can do so when it becomes a little bit more complex I switch to cloud llms and as you can see it it generated something in 12 seconds let's see what it generated so now we have an h calculator application and um it made a mistake interesting it also I I will get to the mistake in a minute so it picked up the the birth date should be part of the URL um as a path variable very good then it parses the date that comes in Via URL very good good and then here's a mistake it should have used today and this works but this is my mistake this is a human mistake I only provided four test cases because I want to keep this presentation short and interesting uh I should have added more edge cases like yesterday today and tomorrow um and if you do that then you get a better solution but then um it's substracts both dates it gets years and it returns this if I would have written this myself it would take me at least 5 to 10 minutes and the tool produced it in like 10 seconds so the business case is pretty simple um the maven plugin works too I am really happy because now we have something that we can use at our daily job so what does work well plain Java code the first few demos were plain Java code and that works really really well because the llm knows about plain Java code the same applies to well-known depend Tendencies I use spring Boot and the llms know about spring boot because there are a lot of examples on the internet so that works also really well what does not work well references to your own code so imagine you have a utils class and you want the llm to come up with something that uses your uus class uh that doesn't currently work because the llms the public Cloud llms do not know about your uus class that you wrote yourself um you can solve that easily and I will talk a little bit about that in a slide called future but currently it doesn't do it also what doesn't work well less known dependencies I use spring boot because it's pretty popular and pretty well known but if you write your own dependency and you use it in your project the llm will not know about it of course so what if it doesn't work we've seen that happening a few times during the presentation first thing you do is retry second thing you do is you check the locks maybe there's a book in the tool or maybe something strange is going on third thing you do you upgrade the large language model you come you go from local to cloud and then what if that also doesn't work then at least you have to unit test no no no hear me out so how did I work before this tool I would write my unit test I would grab a cup of coffee come back and write the implementation nowadays I've write a unit test run the tool grab my cup of coffee and come back hey and if it works it's really nice and if it doesn't work you didn't lose a lot of time so I think that is not a bad thing just spending a minute on running a tool so pricing pretty non-existent when you're using a local large language model not entirely true uh when I'm using those local models uh my battery dies a little bit sooner Therefore your electric bill will be a little bit higher but you probably will not notice it you pay a small price when you're playing around with online AIS like I'm currently doing so building this tool practicing for this talk cost me about $11 so I think that's doable and it's pretty pricey when you start using online AIS a lot let's say you use this for every test case that you have then it will cost a lot of money but on the other hand every time when it succeeds uh it saves you a lot of money because you are the most important and most uh how do you say that the most expensive part of the project humans are the most expensive part of our equation so every time when it comes up with a solution and you say save a lot of money especially while models are getting better cheaper and faster every week so I really believe in the near future we will have a point in which this is cost effective can I use this myself yes you can so I will share four links uh at the end of this presentation a link to those slides these slides a link to the project with the examples a link to the repo on GitHub that builds the jarf fall and the link to the repo that build the M plugin little word of warning it is proof of concept level um like I told you in the beginning I wanted to come up with something really cool in three to five evenings didn't work out took far much more time so this is still very much proof of concept level so be warned future so I re see a lot of things I can improve in the near future to make this more interesting tool and the first is provide multiple test files so currently you can provide only one test file and that's not re realistic normally you have multiple test files and adding that shouldn't be that hard I just need to find the time provide context files I told you whenever you have a utils class it doesn't know about it if you would provide your entire context of your project and send it to the large language model it would know about your uus file and therefore would use it um also small fix now small fix it takes some time but it's not I I can see it in my mind how it should work currently the tool only accepts if the llm comes up with code for exactly one class but sometimes it comes up with multiple classes tool cannot handle that sometimes it comes up with partial implementations especially when you're going through the upper Loop so it comes up with the first implementation and then they say this test didn't work it says oh replace this method with that method method the tool currently cannot handle that um hybrid mode I really love to work together with AI so you as a human write the first version of the Imp mation and you leave todos for the AI implement this implement this implement this currently I support o Lama jet GPT and CLA but there are far more interesting far more AIS out there that I would like to use currently it only supports junit 5 of course you can easily add junit 4 uh but may maybe you want to use Spock or something else um so I would like to support more then currently it only works for Java I think it's easy to add codlin uh but the principle holds for every language so it should work for every language my end goal is to make this an IDE plugin so picture this you're writing your unit test you right click in your IDE generate implementation you grab your cup of coffee and you come back would be perfect I specifically said my endpoint because I think there's another endpoint for this idea and that is part of co-pilot AI assistant of defox Genie I think this should end up in one of those tools actually Stefan asked me if I if I could make it part of deogi and I really would love to but I don't have the time so I have a question for the audience if you have the time um pick this up and probably next year you can stand here big chance big chance um but if you if you if there are people in the room who are willing to make this part of Dev genie please ping uh Stephan of ping me and we will make this happen uh conclusion tdd and gen is a perfect pairing I've shown multiple demos um in which I showed off the fact that the principle Works uh and yeah you can just use it you can play with it um the most important thing that I can teach you today is that you have to check the output of large language models I've given you the example of Air Canada um lucky for us we have a way to check the out of the large language models and that is using unit test and that's what this idea does thank you for being here if you have questions please come to me after the presentation down here I prefer my questions oneon-one um feedback is more than welcome please rate my talk in the app uh if you just have feedback you can come over here talk to me in the hallway um notice that there are QR codes in the bottom those are the four links I promised and last mandatory Bullet From the HR department we are hiring if you're looking for a cool new job talk to me thank you I hope everybody learned something new enjoy the remainder of the conference bye-bye all systems are nominal initialize Genesis sequence
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.