Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Krish Naik · @krishnaik06
Words
21,101
Runtime
4:51:00
Speaking pace
73wpm
Reading time
88min
73 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello all, my name is Krishna Nayak and welcome to my YouTube channel. So guys, this video is about an amazing end-to-end project that we are going to specifically implement and from this particular project you'll get
37 words, the words spoken in the first 30 seconds at 73 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1,396 |
| Average words per sentence | 15.1 |
| Longest sentence | 301 words |
| Questions asked | 64 |
| Sentences containing a number | 97 |
Most used terms
Filler phrases
209 in total: uh 103 · like 41 · kind of 22 · basically 13 · right? 10 · actually 7 · you know 6 · um 4 · literally 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello all, my name is Krishna Nayak and welcome to my YouTube channel. So guys, this video is about an amazing end-to-end project that we are going to specifically implement and from this particular project you'll get a clearcut idea about how a AI forward deployment engineer basically do what they actually do in companies how they implement projects and many more things. Yes, this project will be for somewhere around 4 hours.
So please make sure that you watch it and in this project we'll go step by step. We'll go everything we we'll discuss everything in detail what problem statement we are actually solving and many more things. So first of all let's talk about the problem statement. So here uh you'll be seeing that I have created two important architecture. One architecture is about a legacy system. This is the traditional setup. Let's say there is a client who specifically having this legacy system.
Okay. and then we are going to probably go ahead and convert this legacy system into an AI powered architecture and this will be specifically built by an AFD. Okay. So here uh if we discuss more about the legacy system you can see that you know we have some data like GPS data, truck sensors. So this is basically of a logistic company. So GPS data, truck sensors, route telemetry. Uh one very important thing guys uh this is a kind of a real world project in the companies.
We have just uh you know taken this specific use case. We are not mentioning any company name or something like that. We are just trying to bring the real world scenario that the client had over here to give you the exact picture because that is what this video should be bringing you out a real world use case. Okay. So here you have the GPS data, you have truck sensors, you have route telemetry. All this information is basically going to the SQL server and finally you have with this BI dashboard and uh here you can basically go ahead and find out information like where are my trucks any risk let me check the dashboard everything like that right so you basically go ahead and analyze this reports and try to get out information right so there are a lot of limitations over here like manual analysis queries static dashboard no integration with external data like weather no access to companies SOP no audit trial for user action so suddenly the CTO of this company has said that now we should definitely move to something called as an AI powered architecture and what I have actually done from my company we have brought one afd person okay and we have given this entire new architecture.
So here you can see from the new architecture all the data sources that is going to the SQL server which used to happen in the legacy that is same but on top of it we bring enterprise database then we bring an AI agent orchestrator then you have you can see that from this agent orchestrator you can ask user can actually use natural language processing show me the trucks near Los Angeles all the information then we are considering weather APIs also live weather API so that we can actually give the accurate information then you can see we have also imple will be implementing vector rag inside this then SQL audit trial enterprise trust because this is really important so that we bring up compliances and transparencies each and everything each and every analytics is basically captured over here then we specifically go ahead and uh you know go ahead with the deployment and here we are going to go ahead and consider AWS and uh from GitHub to GitHub actions deploying it to the ENC2 instance and then the live application and finally you'll be able to see the end user output how it basically basically needs to be seen right so the key improvements with respect to the AI FD approach is that we are bringing natural language interface here we are definitely going to use AI models LLM models real-time insights with external data compliance awareness answers using rag autonomous reasoning and decision support full audit trial for trust and transparency automated deployment to CI/CD so from data to intelligence to real business impact that is what a specific AI FD engineer basically do again uh uh this entire project has been recorded by one of our mentors in Krishna Academy that is Monil Kumar.
Amazingly, he has explained step by step. I've also given the timestamp in the description. You can go ahead watch it out. Go ahead and implement along with all the implementation that is basically done. Okay. And uh uh before I go ahead guys, I quickly want to announce that we are soon launching a AI forward deployment engineer boot camp that is from August 4, 2026. The classes will be for on Saturday and Sunday 8:00 a.m. to 11:00 a.m.
IST. This boot camp is for five months uh you know and uh if you really want to go ahead just go ahead and check out this particular boot camp detailed entire syllabus is basically given over here what all things how we'll going to cover and after every modules we will be covering multiple projects so that your understanding becomes very very well right and who is the mentor of this particular uh batch me and Monal Kumar so mon you'll be able to see this specific project from him and then enjoy the project implement together and definitely do check it out We are providing initial 10% off in all the specific courses.
You can just use the coupon Christian. So yes uh let's go ahead and enjoy this specific project. Implement it along with the video that we have done. Follow the instructor and try to execute till the end. Write some amazing comments in the description. At least we'll keep a like target of thousand. So please go ahead and enjoy this entire product implementation. Hey everyone, I hope you all are doing well. My name is Monal and in this video we are building an end toend project but from the perspective of an AI forward deployed engineer.
See most tutorials teach you how to build models, how to create AI models, create AI workflows in a perfectly clean environment but that's not how industry works. In the real world, you are usually handed messy legacy system, old system, old databases, strict safety around data. It means you require someone else to get the data for you. And clients who don't care about your text stack, it means you are using lang chain, you're using langraph or crewi, they don't care.
They only care about whether you are able to provide them business solution or not. And this is exactly what we are going to see in this particular project. Legacy integration, documentation, business presentation, system safety and client ready architecture. A little disclaimer before getting started that I will try my best to deliver you this project as simple as I can. But this project is intense. It is complex and it is exactly what will get you hired.
And without wasting any more time, let's dive in and start this awesome project. So I'm calling this project as cold chain logistic and we are building here AI assistant. And before talking about what this project is, what is cold chain logistic, I just want to give you clarity on the prerequisite because if you don't know all of this prerequisite present here, you are going to have difficult time understanding this video.
But if you're someone who understands this, keep this in mind that I'm not going to teach you what is an LLM, what is a vector database, what is a langraph. I will be applying them. So this is more like a applied concept. We are not going to learn any concept. We are just going to utilize that concept in the project to create that business solution. And if you are someone who does not know any of this tools or any of the skills or any of this prerequisite then it is still fine to attend and to watch this video to understand more about how a seasoned professional like an AI FDE approaches to a particular solution.
All right. So these are the prerequisite and let me state all of them. So for core programming and data fundamentals we are using Python as the main backend programming language or the main core logic language. Then obviously we will require some object-oriented programming that is the standard for any language and then we will be using some SQL database. I will talk more about what SQL database exactly we are using but if you know the language of SQL that is more than enough.
All right later thing I will be teaching you. So SQL we require here in terms of artificial intelligence and LLM core we are required to understand how to call an LLM API call how to call an LLM what are API keys how do we set up that particular path and if you are someone who knows lang in lang graph then that makes it very easy to call your choice of llm then vector database I will talk about what kind of vector database we are using but if you are familiar with an idea of what is a vector database why do we use it then your prerequisite for this video is already completed.
Now if I talk about advanced AI agents and frameworks we are going to use langraph we are going to use AI agent we are going to create AI agent so we require that agent memory etc. That's why I have written that you should know what AI agent is and tools what are the tools what we call tools what AI agent call as tools right this is nothing but function but function that AI can call but again I'm expecting this that you guys know that second thing is last thing sorry about that that development environment and tools so I will be using VS code for some reason I will talk about that and you guys can use venv or cond I'm personally a fan of cond because I find that that is more stable but we can use Venv.
So for my system development I will be using Gonda but when we are going to deploy this project there we will be using VMV. Okay, Python virtual environments. Final thing is user interface and deployment. So we are going to use streamllet because we are not front- end engineers but we are going to showcase this project as a P or as the final thing that business can use. The stream we are going to use it for that UI application and docker desktop should be installed in your system along with you should know or you should be little familiar with AWS things like what is security group what is IM I'm not going to teach you how to create AWS account instead I will be using the AWS account that I already have and we'll be showcasing that how to create that EC2 instance what are security groups how to get started etc.
All right. So these are the prerequisite for this particular project called as cold chain logistic AI assistant. Now let's talk about the problem statement. Let's suppose there is this logistic company that stores their supply chain data in one of their legacy system over cloud. And the problem statement is they want to talk to their data. That's all. I know this is pretty vague statement but that is usually what an AIFD engineer gets and from there what we usually do is we talk to the client we understand what is the exact problem that they are facing what are the things that they already have where is the data can we get access to the data what is the solution you are looking for so we are going to propose them a solution if they like it we will proceed with the coding if they don't like it we will propose them a new solution what are the different security benchmarks that they already have.
We are going to gather all of this scenarios before starting the actual coding. And that is the main difference between a usual AI engineer versus a AI FD engineer. So let's talk about the problem statement a bit. I'm going to write a very vague statement because the problem statement itself is very vague. So we have a client and client have their data in the cloud. So let me create that box here that is going to tell that this data that is present here is present in the cloud system.
All right. And what client wants is client wants to talk to their data. So in short client wants to chat with their data. Okay. Chat with the data. This is what they want. Now this data is present in the legacy system. So this is not just any data that you can uh dump in your system using API. So this is a legacy system. And when we ask clients what are the requirements they will be giving you that yes there is security behind it but we will be giving you an ID and password because we don't maintain a separate uh DBA in our system.
Sorry the password spelling is wrong. So let me write password. So they will be providing us with the ID and password and that is the legacy system that they have here. And let me talk more about this legacy system. So this particular database is a relational database. It is MSQL data set and yes we don't have access to the API. All right, we will not get their API. It means we cannot dump the data. So it means we need to get inside the database and proceed with the solution.
We need to do something on this particular part. And usually whenever a client say that they want to chat with their data or talk with their data, it basically means that they already have some means to talk with their data but not in the natural language sense. So what are the different ways that we can chat with our data? So now let me expand this further. So one of the ways to chat with your data is dashboard. All right.
And dashboard is not that magical software that answers your query. It you need to learn how to use the dashboard. And behind the dashboard, usually there are developers who create multiple API for the dashboard to consume. And this API is directly connected to the database or someone needs to create the back end for those APIs. All right. So once the API is created, then user can use the dashboard. But what's the problem with the dashboard?
Whenever we say dashboard it means there needs to be that UI and in that UI there is going to be lot of lot of drop-down menu all right there is going to be lot of drop-down there is a learning curve learning curve we need to have charts and behind all of that let's suppose we are talking about the logistic data here there are some SOPs so people are required or business is required try to understand those SOPs before even getting started.
So there is a learning curve before this dashboard and there are a lot of things. Let's suppose I want to talk about the data specifically supply chain data for the year 2024 January from the uh from the date 01 till the 15th of January. to do this specific thing in the dashboard. They need to know exactly where these values get filled, what kind of graph they are looking for or what kind of analytics they are looking for.
If they don't know this, then the dashboard is going to show you hundreds of different answers or metrics and you are required to understand that metric. So that is a lot of learning curve that as I already mentioned before and when they say that they want to chat with their data this is one of the way and the other way is have that agentic AI solution or a chatbot based solution. Now why this is more better as compared to dashboard?
Yes, dashboard is going to always give you accurate information with API and database. So your agentic AI solution can give if you know how to deal with that whole engineering that we're going to learn in this video. All right. So when we have this agentic AI based chatbot usually what happens people can naturally talk that I want to get the supply chain logistic data for year 2024 month January and date between 01 that is 1st till 15th of January. give me the analytics and you can directly ask that I'm looking for so so so matric and when you pass that your chatbot or your agentic solution it is going to go into the database trying to retrieve the data and then once it has the data it will try to make sense of the data and provide you the final answer.
So here usually we have much lesser steps as compared to dashboard and the learning curve is you just need to know the language you just need to know what you're looking for from that AI powered chatbot. All right. And that is why business usually tend towards this chat with the data. And for this particular client, the imaginary client that is logistic company that stores their supply chain data in their legacy system, we need to create we need to cater them a solution so that they will be able to chat with the data.
I will talk more about this legacy system where it is. But let me talk about what things I already have here. uh first database where is this data we are not just going to select any data a CSV file in the Python create from the notebooks no when we say legacy system when we say that there the data is sitting behind a security it means we need to connect to the database so I've already prepared I have already prepared the legacy system database that is running in my Amazon EC2 system okay so it is running in my EC2 system I will talk about the steps that I did but that is just a script file that I will be giving you which you guys can run but for this whole video we will be considering that this database is already running we just need to connect to the database and do the later process all right so this is already a running instance in my system let me show you that this is the running instance not this one but I have it here this is currently not running but let me start this this is cold storage logistic mysql so I just need to click on it this is my Amazon EC2 instance page and I will be starting the instance.
All right. So for this video my now my instance is running and I will be getting that public IP address right now because it is not starting it is not giving that but soon it will be starting and we are going to copy and paste the same thing in the proceding sections of this video. All right so now that our database is ready the next step is once we understand the problem statement we need to create some kind of solution.
Okay. So once we have the problem statement, now that we know the problem statement, it's time to decide the solution. So we will be doing solution discussion in upcoming videos. The solution discussion and then create a presentation for the business because this is the solution that we are thinking but until we are able to communicate the same to the business they will not be able to take the decision. So that is also a very important part for an AFD engineer not just doing the engineering work but also doing the client interaction also doing the client handling.
So we are going to have this problem statement that is already known now. So let me write that as known. And then in the next videos in the next section okay we are going to discuss about the data what this data is and how do we create solution and then create this presentation and then the coding part. All right. But the problem statement is we have the client that is having the logistic supply chain data in their legacy database and we just need to chat with the data.
As you all know that for this particular project we are considering a imaginary logistic company that is saving their data in one of their legacy system over cloud. But that does not mean that we really have the legacy system. But we are not going to fake a legacy system with a JSON file or a CSV file present locally in our system. What I did, I took this particular data which I will talk about in a bit. I ingested that data by creating a EC2 instance starting a MSSQL 2022 server there and ingested the data from my local system to there.
So the data is present over the cloud. I will be teaching you how I ingested the data there. You can start your own server and then we can start the project. But I already have that running. You are not going to get access to to that public uh EC2 instance because I will be stopping that after this project. But you can create your own for the practical part of this particular project. All right. So what I have here is that I'm using this particular data set that is logistic and supply data set.
So let's understand more about this data set and we will talk about the injection part and other parts later about the data set. Let me zoom in so we can see the text clearly. So this data set captures a comprehensive set of logistic and supply chain operations specifically collected from a logistic network in Southern California. All right. So the data spans from January 2021 to 2024 encompassing various aspects of transportation, warehouse management, route planning and realtime monitoring.
Okay, this is the actual real-time data for those logistics for those transports. Okay. It includes detailed hourly records of logistic activities reflecting conditions in urban areas and transport corridors known for high traffic and dynamic operational challenges. You can see that the usability of this data is 4.71. That's why I have considered this. So this is the data. This is the back end or the core data of our imaginary logistic company.
And you can see this data set is collected from various sources such as GPS tracking system, IoT sensors, warehouse management systems, external data providers, etc. Now let's talk about what are the features that we have for this data set. And when we say features, don't worry that I'm not going to create any machine learning model here. Feature means just the column name. So I'm having timestamp for that particular uh row of data and then vehicle GPS, latitude and longitude.
What is the fuel consumption rate? What is the ETA in hours? The difference between estimated and actual arrival time. So we have the ETA. Then the traffic congestion level. Warehouse inventory level. Loading unloading time. Handling equipment availability. Okay. Like forklifts unavailable or available. Order fulfillment status that it is fulfilled or not fulfilled. So when they deliver it, it will be fulfilled. Weather condition severity, port congestion level, shipping cost, supplier reliability, lead time, the average time taken for supplier to deliver materials, historic demand, IoT temperature, the temperature by IoT sensors in terms of degrees C.
So this is important information to keep in your mind that the temperature that is present there is in degrees CC. Then cargo condition status whether the status is poor or good. So a lot of the times when we talk about this um logistics right the temperature matters a lot and it matters more when we talk about when we are delivering something uh for healthcare right so route risk level custom clearance time driver behavior score fatigue monitoring score target variables we have we we don't require target variable here but there are a lot of things like disruption likelihood okay a score predicting the likelihood of a disruption occurring so we already have that value delay probability what is the probability of getting delayed that is also present in the data.
So we have multiple predicted or you can say dependent variable or target variable based on all of these features. All right we have delay probability we have risk classification a category classification indicating the level of risk. So we can also utilize this delivery time deviation and this data set can be used for various application in logistic supply chain management including you can check the risk assessment and disruption detection optimization of routing and scheduleuling to minimize delays predictive maintenance for logistic vehicles and analysis of impact of external factors such as traffic weather on delivery times.
So here we can also use this data set that to predict the factors whether it will be delivered or not and whether the factors like whether end traffic has any effects on the delivery time we can also check it. So what we are going to consider we are going to consider that this is the live data from the company. So 2024 is going to be the latest year for this particular project and we can query this database. We can understand we can ask things like how many uh logistics or how many you can say tell me how many different vehicle details are present in the database.
What are the risk associated and how many can you can you guide me on the different types of risk that we have in this 15 days or first 15 days of January 2024. We can ask lot of different weather severity conditions that based on this weather conditions how many vehicles do you think is not going to pass. All right. So once we have decision like this then business can take action on it and can use this chatbot directly without using the dashboard.
Okay. And because this is a AI chatbot we can create a SOP for the business. Right now the SOP is not available here but I will also be creating a simple documentation which based on it our AI chatbot can take certain action. All right. So what we have here is that in short this data set contains vehicle details plus risk plus traffic congestion plus uh deliver deviation right the ETA the weather condition severity etc.
So we are going to use this data set we are going to ingest this data set. So we have a base or we have a legacy system in place. So once I have this data set what I'm going to do I'm going to download this particular data set and once we download it we need to keep this in one of the folder. So what I'm going to do to track everything I'm going to create a folder structure here initialize the git and then going to push this in my GitHub that is present here.
So I've already created a repository name cold chain logistic FD projects and I'm calling this as cold chain because here we talking about temperature conditions right based on the temperature conditions if temperature is below this what is the severity what is the risk level for this particular transport for this particular logistic. So we are going to create our project around that. All right. So let me create this repository name and I'm going to keep this as public and the link to this data set I'm also going to write it here.
Okay. So first let's create it and then initialize few thing in my repository. This is my VS code. I'm using VS code as our hybrid IDE and I'm using Windows. So if you're in Mac please adjust accordingly. So what I'm going to do first I already have my cond environment ready but I will be giving you instructions for the pip. All right. So first let me create a folder called as data and inside this data I will be keeping the data set that we have.
All right. So I will be creating a folder called as raw and inside this data I will be pasting that raw. Okay. So this is the same data set. And this is the same data set that we saw here. And inside this data set, we have the time stamp, vehicle, GPS, latitude, longitude, etc. present here. And all of that are numbers. So it is good for us. Okay. Not everything is number. We are not creating machine learning. So we don't have to do any feature engineering.
Our L&M can understand all of this moderate risk better than just encoded labels. Okay. So I have this data set and also let me create the data link. Okay. What is the source? What is the data link? So for the source, what I'm doing, I will be creating a folder called as source. All right. I think I created this here. Yes, I want to move this into data. Let me zoom in so you all can see better. And then inside this source, I will be creating data.txt.
So everything is noted here. Okay, let me go there and copy the link. So you guys know where to get this data. Okay. All right. So we have this data here and in raw we already have this data. So what I'm going to do I'm going to use this data to ingest in the EC2 instance. I will not be talking about EC2 instance much but I will be giving you some instructions. So let me give you my instructions here. So I will be creating docs.
Okay. And inside it we will be maintaining our file that is for us. You guys can use this as the context for this particular project what instructions we are following. So instructions do MD and inside it we'll be writing. So the step one is let me not not write steps but we are talking about injection or ingesting data. Okayesting data. So first step is download the data set from let's refer to this particular source.
Maintain the documentation from day one is a good sign of a great engineer. Okay. So we need to document as much things as possible. Now the next thing is create an EC2 instance. Create an EC2 instance and we need to export some ports. Okay. If you are not familiar with how to create EC2 instance, I will suggest you to watch some videos on that. But I will be showing you step by step how to create EC2 instance for the streaml UI that we are going to do later in this particular video.
But what I did I used I created what I did is that I created a docker container. Okay. For my SQL server this one. Let me show you that. Yeah. So this is the docker container that I'm using. Okay. And I'm creating that image. I'm running that in the docker container present here. So as you are able to see this is running. And now if I click on it, you'll be able to get the public IP of this one. So what I'm going to do for now, I'm going to keep this in my environment.
So we can connect to this later. There are a lot of configuration that I'm doing right now, but don't worry about it. All of this will be simplified later once we are little comfortable with all this information. Okay. So this is the database credential that I'm writing that my database setting is going to be my SQL server host. Okay. Is going to be this particular URL that we able to see. So this is public URL. And then I am running this particular port.
I'm running this particular container in SQL server port of 1433 that is a by default of uh MSQL Microsoft SQL server. Okay. So 1433. So now let me talk about how to even create this instance. First you guys can try this in your local system but for that you need to have docker in your system. Docker desktop. In my system docker desktop is running. So if I do docker ps a you will see that my docker desktop is not running.
So first let's start the docker desktop. Not from here. Docker desktop. It will take some time to load. All right. Starting the docker engine. Okay. So you can see this is the docker that is already running not running. So we can see docker ps a to list all the containers that we have created or I have created. You can see that I have created this particular container. It is running in the port of 1433 and this is the TCP connection that I did.
And this is the name of the docker. How you will be creating that? Let me show you that particular instruction. So you can use this instructions to spin up docker containers and the choice of your environment. I have given you multiple code here. Let me remove this instruction from here. So this is to spin up the legacy MSSQL server in your Windows system. This is the single line docker command. This is multi-line in Windows.
This is multi-line in Unix. And this is multi-line with volume inside EC2. Because when you're working with EC2 instance, we will be starting and stopping the EC2 instance because we don't want to get uh unnecessary billing on that. So if we stop it, whenever we started, we want to get the same data present there. All right. So I will be just showcasing you what I did to create that instance and ingest the data there.
But I already have that instance running which I can connect to and can retrieve the data from there. So let's start with the spin up legacy MSSQL server. So what things we have here? So when I say docker run hyphen this is the configurations of the image itself while we are creating the container. This is specific to this particular container ms mcr microsoft.com mssql server 2022 latest and I'm using this particular configuration called as accept ula equals to y which is nothing but terms and condition.
So msql requires you to accept the terms and condition otherwise it will not be even starting. This is the first configuration. Second configuration is the uh password the ID and the password that you require that is the first user of this particular database and I'm mapping the port for port 1433 from my system to the docker system and the name of this particular container is going to be legacy msql fnd is nothing but detached mode.
It means whenever I'm going to run this particular command it is not going to uh it is it is going to free up my terminal so that I can use another command. Okay, this is detach mode and this is the name of the image that I'm using to create a container because this is already present in my system. I can show you a few things. If I write docker images, you can see that this particular one is already present in my system.
It is consider it is taking like 2.34 GB of my system and it is only 625 MB. Okay. So I don't have to uh run this particular command or pull the image. I if I just run this, it will start a container. But let me first check if a container is already present in my system. Yes. So let me remove this part because I don't need it. I will be showing you things from the scratch. So docker remove this. Okay. Sorry about that.
So docker remove is not the command but the command is first let me stop the container. if it is stopped then then the docker rm and then the container id. So rm is for container and rmi is for images. Okay, I don't want to remove the image. So I'm just removing the container. And if I do docker ps- a again now you will see there is no container running in my system. Okay. So now let's run this particular command. And again I'm repeating this is the same exact command but when we are working with EC2 instance we need to run this one more thing that is the volume.
So we are appending the volume. So whenever we save something it will be saved in this particular in this particular direction in this particular directory in the Linux or Unix system. Okay. So this is something that we need to add or keep this in mind while working with EC2. But for now let's focus on this particular system. Okay. So this is docken run. Yes, we accept it and let's run it. Okay, so as you can see it is not downloading anything and the system is up and running.
It is that simple. Next thing we have to do is we need to ingest the data and for ingesting the data I already have the script ready. We don't have to uh create the script again. I will be creating the other code step line by line but for the injection code we don't need to create anything. I will be creating scripts folder here. We will be using the generic script for this particular project and one of the script is ingest legacy data. py.
Okay. Ingest legacy data. py. So what it is going to do? What we have to do is first of all we need to connect to this particular data uh database okay docker container and then whatever the con instructions of the CSV is present we need to load it and then dump that in the database and once that is done we have that legacy system that client usually have we going to do the same thing for EC2 instance I already have it running but this will give you an idea of how to set it up okay so let me do that particular part uh Let me bring that code here.
Okay. So this is the code for ingesting the data that I have used. So what this particular code is saying that we need to have SQL alchemy create engine etc. So before doing any of that let me bring you the requirement.txt for this particular project. Okay. So let me bring that here. Okay. So these are the requirement.txt. Elier pass CPU lang chain community langen hugging space langchen openai pine cones we will be using the the pine cone vector database lang graph checkpoint lang graph lang smmith open aai open spy excel to read the excel file I will talk about that again later but I will suggest you to set up this particular environment or the environment that you particular like but I already have the cond environment so I'll be using that so let me first check env list and you can I've also written the Python version that I have used in my system for this particular project and I will also suggest you because all these versions is with respect to Python 3.12.13 okay and I am using yes FDE env test so let me activate that activate this particular environment all right so now that the environment is activated I don't require any installation but if you still require it so you can do pip install -r requirements txt.
Let me also add this in the instructions. Okay, we don't require it. So let me clear it. As you can see my all the instructions are successful because they were already satisfied before. Okay. So in the instructions what we have to do before doing after doing a docker run is that install the requirements. Okay. So I'm making it easy for you to follow. Once the docker container is up and running then use this command pip install r requirement.txt.
And before moving any forward what we can do is we can initialize all the g commands that this repository is saying that is get in it get add all the files get commit and get branch m okay let me copy this part and others I will write it quickly okay so first let's initialize the g then get add dot and get commit what we are doing get commit- M adding the instructions and injection file. All right. And then setting the branch to main.
And then setting the remote URL. and then finally pushing it. All right. So whenever we make some big changes, we are going to commit it. We going to push it into GitHub repository. So even if we lose something in this particular video, we have all of them tracked. Okay. And one important thing is that we have pushed the env. Right now there is nothing in the envir. It's a good time to have the get ignore file. Okay.
So get ignore file is going to restrict few files to be ever committed and that is what we want and you guys can create virtual environment but if you create virtual environment and py cache let me add all we have added py cache vnv.vnv etc. And let me save this and let me commit all this again. So get add dot get commit. Let me add added get ignore file. So our env is still secure because we didn't have anything there.
Let me commit that and get push origin map. Okay. But for some reason I'm not able to see this one getting reflected here this env. So just give me a second to check if it is getting pushed there again or not. So where is our GitHub repository? Sometimes it does not get reflected in VS code directly. So env is still present which should not be here. So get ignore this file looks good. Nothing else is here. So as you can see that it is still pushed.
So what I'm going to do I'm going to remove it from the caching and once it is removed then what we will do we will try to add all the things again. So hopefully it will work this time. So let me get add dot but you can see that VS code is still not reflecting this particular part. So let's try get commit- m or I can write the same commit command that we have used before. Where is that? Yeah. Okay. You can see it is now in deleted mode.
And let's push to origin and let's see what happens. Okay. Hopefully we will see it updated there. and we don't see any file. Okay, now the env file is not getting tracked anymore and we can update it without having to worry about it. Okay, so get ignore file is also updated. We have all the codes. Now let's talk about the scripts that is a injection script. So what I'm doing in this injection script, I am using pandas to read the CSV file and using the pandas system to directly update or push that into the database.
But there are a few more things that we have to worry about that is let me show you that that is a version of database present there that is a driver that is present inside the docker container. Okay. So once that this is ready let's activate few things let's install few things and let's try to fix one thing. Okay. So give me a second. I'm looking for that particular version which will show you what is not correct. Okay.
So it is not going to work in my system anymore because I've already fixed it. But first let's focus on this code that run it and then once we receive any error or if you receive any error I will guide you how to proceed with that. So let's talk about this particular script. So what it is doing I'm currently reading the parent of this particular script that is nothing but scripts folder. Then I'm checking what is the parent of the scripts folder. that is nothing but this cold chain logistic folder name.
So once we get the project root then I'm locating the env from that env we will be loading few instructions we'll be loading few things from there. Okay but for now what I'm doing from the project root I'm locating the CSV file where we have the original data and from the env we need to load few env files which we don't have it yet. So what I'm going to do I'm going to create few uh system admin. We already have what is the SQL server host.
If not the if the host is not present then it will be looking into the local host. But if the host is present then it will be utilizing that URL that URL of EC2 instance. Okay. So what we require is we require SQL admin user as well as SQL admin password which is not present here. So let me write the credentials that I will be using for the database. The credentials are SQL admin user sim and I'm using the same password which we used while creating the container.
These two things needs to match with the container username and the password. Now you might be asking that where we have mentioned the user SA. Well, it is present in the configuration itself. So SQL server or Microsoft SQL server is using a particular pattern to check the password or the user. The first thing is the database name itself. Then the essay is the username. Then the password means we are setting the password and what is the password.
So within this particular string it contains multiple entities and that's how it knows the username is SA which we are writing it here. So user is SA and the password is same as what we have written here. So what we require? We require SQL admin user, SQL admin password, SQL server host, SQL server port which we already have it here. So if I just comment it, if the env is not found, then what will happen in the legacy while ingesting it?
It is going to select the local host and we are going to do the same thing. So we want to select the local host first to try it out how to uh start this particular server and how to connect to it as well. So once this part is done what we are doing we are reading the CSV and then we are mapping it to the enterprise schema. So what I'm doing we have few things like time stamp vehicle GPS latitude etc. And to create a legacy system and to get the sense of a legacy system we are creating weird column names or column names that does not tell anything just to showcase whether this is a legacy system or not.
For time stamp I'm calling ts UTC vehicle GPS I'm calling V latitude vehicle latitude vehicle longitude IoT temperature I temperature value in Celsius. So I'm creating that proper MSSQL database column names that uh is used in the legacy system that is not just a pandas data frame. So I'm mapping the actual CSV data that is present here. These are the actual names to the names that is usually traditionally used in the legacy system.
All right. So once I have created all those mappings that what I'm doing I'm extracting what are the different columns there. Then I'm adding a new column. This is not at all required but if you want this will tell us that yes these are the columns that is ingested just to have a better sense of the data. Now this is the part where we need to connect to the legacy database to ingest the data. Usually this whole process that we are doing in this particular section, this is already done.
Business will give us the instructions to connect to this database. Business will give us all these configurations and we will directly be connecting. But because we are learning here, I will be showing you how do we even create this whole thing. So we can replicate that in our system. But you can skip this total part of injection if you don't want to worry about creating your own legacy system for this particular project.
Once this is done we are creating because this contains lot of string as well and it requires in such format we are passing this particular connection string and passing that to the final parameters of MSSQL pi ODC ODC connect when passing this to the parameter it will create that engine using that we can insert the database now once we have this engine what I'm doing pandas can directly ingest or dump the data in the database so this is the data frame df legacy we are doing dot2 to SQL table name we need to provide the engine engine contains the connection to it and if the data is already present there replace it if data is not present index we are not setting any index and we are creating a schema of dbo okay this is the main schema that we have and once this is running once this is completed it will say legacy system injection completed so the next thing that we have to do is let me copy the relative path from here and let's write the same thing in the documentation instructions that once we install the requirement the next thing is we need to do python scripts in legacy data py in this my in my windows I might need to write with this but let's see okay all right so let's try all right so this is working it is ingesting into the double and then the legacy data ingestion completed So if you're someone who does not want to work with the EC2 instance and create and you don't have free credits then don't worry about it.
This particular project is totally compatible with local. So if you want to build it you can build it from uh scratch by watching this video. Now let's suppose this local container that is running in my system is the business legacy system. So we have the business legacy system running in local. We also have or I also have the business legacy system data running in my EC2 instance. Now we are in the same stage. Okay, we all are in the same u area.
Okay, what we can do from this one is how to connect to this database. The one thing the one approach that I personally like is connect via VS code. I have chosen VS code specifically because the extensions that are present here there is one extension called as SQL. Let me show you that instruction. Just give me a second. So what we have to do is that there is something called as Azure Studio. So you guys can also download Azure Studio.
Let me show you that. So let me give you this instruction. Okay. All right. The next instruction is connecting to the data. Data injection is completed. Next is connecting to the data. You can follow the first step which I'm not following. I will also tell you why. But if you go to this particular URL that is Azure data studio, they will also recommend you to use Visual Studio. You can see the content may have been retired and may not be updated in the future.
And they're also telling us that you should migrate to VS uh code, Visual Studio Code. That is what we are also going to do. Okay, we are going to use Visual Studio and we are going to use this extension called as MSSQL which I already have. So if you search MS SQL, you will get the extension from the Microsoft which you can use directly to connect to your MSQL database. So you can install this particular part which I have already installed.
So let me also write that in the instruction. The my recommendation is to use this. All right. So this is my recommendation SQL server to use MSQL by Microsoft and if you use this particular extension you will be able to connect. Now that our database is running in a container there will be this icon not this one. When you install this you will get multiple things like database project and SQL server. You are able to see on the left side of my screen.
We are expected to use SQL server here. Okay. So click on this SQL server and you can see there are multiple connections. It is currently loading. I already have multiple connections. So what I'm going to do I will be creating more such connection. Okay. So let me create let me delete this particular connection here. So I can show you delete. See not the container I want to remove the connection. Yes. So let me create a new connection.
I will also be writing all the instructions of those connection here. But let me first start by creating this new one. Okay. So let's add a new connection and we can write some profile name. So I'm adding a profile name called as legacy MySQL and then server name is going to be there is this is not a server name. This is good. We are using the local host. If you want to connect to your EC2 instance then you will be writing the public URL the public IP address for that particular EC2 instance and you'll be able to connect and you can see it is already by default to 1433 as I already mentioned.
Yes, we need to trust the certificate otherwise we will get some error or it might create some issue later. And we are fine with SQL login. We need to provide the database username and password. Whatever the password you have created while creating the container, use it. I want to save the password and database name is master. That is a by default. Okay. Even if you don't select it, it will connect to master. But here we want to be more precise.
So I'm selecting master here. And encrypt mandatory. I don't want to encrypt it. So I just want to make it as optional. Okay, you can test the connection before connecting. It is saying that yes connection succeeded and let's connect to this particular data set. Okay, so still we are in this process to how to connect to the database. Later we will talk about the security agent and other stuff. We are still in legacy system phase which is very important part.
Okay. So once this is connected, you can see this is connected and there are few things that I wanted to teach you here that once this is connected. Okay, let me also write all the instruction in case if you need it. So this is from Microsoft. This is the download part by SQL and that. Okay. So I've also given all the things that I personally use that is profile name, local host in case you forget it. And once this is done, now it's time to finally connect it.
If you want to run any SQL query here, what you what I will suggest is create a new temporary file in the memory. And to do that, please press ctrl + n in VS code, which will give you a new file. You can see this is creating a new file. But this is currently a plain text. And to run any SQL query here, what you need to do is click on this plain text. If you click on this plain text, it will ask you to what kind of language this particular file you want.
What file we want here is that we want SQL file. So let me write SQL. And once you click on the SQL then whatever code you will write here it will be rendered as SQL. Okay. So what we can do is we can select we can run some particular query here. Let me run this particular query which I'm also going to write. So after this, after the instructions, select SQL and once we do that, paste this particular command to test it.
And right now you can directly click on this. But you need to check whether this particular file is connected to your SQL container. To check it, what you need to do is you can see there is this new icon called as MSSQL. Click on it. Not this one. Local host. Okay. So we are able to see this local host. Click on this local host. Which one it is connected to? We want to connect it to legacy MSQL which we just now created.
Click on this. And now if you run this file control A to select the query and click on this execute query and this query will be executed. So this query is particularly right now saying that it has total number of 32,000 uh and 65 number of rows present here. It means our connection to the legacy database is completed. Okay. to do the same thing here. What I can do, I can create a new connection here. This time I will be calling it as let me follow the instructions.
I will be connecting it to my cloud legacy MSSQL. You don't have to do this. I already have this ready. And instead of saying local host, what I will be doing is that I have my env. I will be copying the URL. All right. And then where is that? Yeah. Okay. I have to create it again. So let me first type the server name. Then let me type that cloud legacy MySQL name. So this is cloud legacy MSSQL. The port name is good.
Authenticate via this. And because I'm using the same instruction to run in the EC2, I will be using the same username and password here. Okay. And then yes, I want to save the password. Select the database. So currently it is fetching what are the different database name present there. So let's wait. Okay. Uh why it will throw an error? Because right now I need to go to my EC2 instance to fetch my live IP address. Okay. because it is pointing to right now a different IP address.
This is something that you need to do. So I have lot of security groups. You guys can create multiple security groups but two security groups we need to worry about. Okay. First is about the app server. Second is about the database server. For now let's focus on the security group called as FD code logistic database server. So create your new security group and then follow whatever the instructions I have present here.
So I'm clicking on the security group and I want to edit some actions. So I want to edit the inbound rules and you can see that I have added multiple inbound rules here. First is I want to be able to SSH to my EC2 instance from this particular IP address which might not be correct. So what I'm going to do I want to be able to access via my IP only and I want to enter to this MySQL using my IP address again. Okay. So there are different types SSH, MySQL.
You guys can also write custom and this one is nothing but I will talk about this later. So there are two configurations inbound rule you need to add while creating the EC2 instance. SSH my IP and this particular whatever your current IP address is my SQL this IP whatever your IP address is and click on save rules. Okay, I have not made any changes. It means this particular one is correct. We don't have to worry about it.
So let's cancel it. Let's go to VS Code and see if it is working. Okay, so for now it is not giving us anything for some reason. All my password looks correct. Let's look at the EC2 instance again if everything is working correctly. Okay, this is our IP address that I'm having. So this is looking correct. Name is also correct. But let me type master here. Encryption is optional and let's click on test connection. So because it is taking so much time in my system there are two possibilities.
Um I wanted to guess but I already have the answer. First thing is we are not able to connect even though our EC2 instance is working. So what can be other issue? The other issue is that because the instance is running it does not means my docker container is automatically running as well. So what we need to do there we need to go to this particular um instance go inside the container and then run the container itself.
So I already have that PM file. So let me go to that particular PM file. Go to that EW AWS and open the terminal here. And then I will be SSH to my SSH I then my PAM file. So I'm using PM file to log into my EC2 instance. And then I will be using the password, the username, the by default name because I'm using Ubuntu. That is going to be Ubuntu. And then the IP address. So IP address is not this one. My IP address is this.
Where is my IP address? So here is my public IP address. Yes, I want to authenticate. And then permission is denied. So something is wrong. Okay, let's try with private one. It is taking some time. Okay. So I can see that literally I have written Ubuntu incorrectly. Let's press enter. Let's go inside it and we are in. So let's check our theory. That is not the theory. That is what it should be happening there. So docker ps a and you can see.
Sorry about that. The PS is without the hyphen and we'll be able to see the container that the container is created 6 days ago and it is exceeded 23 hours ago. So what we need to do is we just need to start the container. So let's start the container and now everything will be fine. Okay. So now that the container is running, we can exit from this particular docker and we will be able to see that VS code will now be able to find the database.
Let's click on test connection. All right. So now we can see that it is finally connected. Test connection is working and if I click on connect then the connection will be successful and we will be able to run the same query which I was running. So let's connect to it and then it is connected. Now instead of selecting the local host I can select the cloud legacy MSSQL here and then try to run the same query to showcase that this is exactly what I have done there.
It is the same number of rows it returned that we were able to see that on the local host. If I try to run the same thing in the local host, you will see that I'm doing the same thing. So this is what I did in the injection part. And if you don't know how to create EC2 instance as I mentioned please I will suggest you to to recall the EC2 instance creation part or wait till I create a EC2 instance for the streaml otherwise feel free to stay in this particular path stay in the sequence to learn more about this project.
Okay. So let me close it. We are able to understand this and we can run few more things about what things are present in the data set. So we can run next thing is this. So let me select it not the local host and legacy MSSQL and let's run it. We are able to see the same data but the name is tsutc vlet v longitude etc is present here and we also added a new thing that is yes this is system ingested flat that we are able to see all right so let's close it and this completes our data injection to the legacy system which I have already created in my EC2 instance so now we have the legacy system that business is talking about the first part of the problem statement that we usually don't have to do but for this particular project we have to do it that is the legacy database.
So this is the first requirement to create any project and we went to the brain uh to create this database. Okay, now it is present. And if you want to know more about my EC2 instance that I created, I can show you a few more things just so you can replicate. So I am selecting my EC2 instance and then I can talk about what are the instance settings that I have chosen. Okay, not this autoscaling part etc. But this one.
So you can see the instance that I am using is where is that? This is the part that I want to show you. Just give me a second as else I will be writing all the parts for this. Okay, this is not showing us the storage. Okay. Yeah. So this is about the storage tags. I'm just checking if I can get details to my SQL. What are the configurations that I'm using? Yes. So I have used the instant type to C7 flex.l large and also let me write all those things in the instructions instructions for EC2 instance and we are talking about database few things I'm using I'm using instance type of C7i flex large and I'm also using a storage of 30 GB GB because we have database here.
Uh we need to run the docker as well. So that will require some space. I'm using storage of 30GB and I'm using Ubuntu. Okay. I'm using Linux based operating system. You will see Amazon as well. But don't I'm not using Amazon. Okay. I'm not using Amazon OS or Amazon provided Linux uh image. I'm using Ubuntu one. Okay. So keep this in mind. Another configuration you can sec but I have also created for this I've also created security group as I mentioned.
So create a security group for this beforehand and attach the security group while creating EC2 instance. So let me quickly show you this particular step here. Okay. So, what we'll be doing if you want to create your own EC2 instance, the first thing that I will suggest you to do is go to security group. Search security group. All right. Click on security group. Then create a new group. Give it a name. You don't require the description, but you can add some inbound rules.
I have added the inbound rules. Custom TCP you can select. But what I did is I should be able to SSH to this particular system. So I want SSH uh and I want to SSH from my IP only and I have added few more rules like MS if you write MSSQL it also is one of those categories that is already available in inbound rules. If you click on it it is automatically a port range of 1433 it is automatically TCP. You guys can add two different rules here.
Okay. Once the rule addition is done create security group. Once the security group is created, you can go and search for EC2 instance. If you click on EC2 instance, what you have to do is you need to click and launch the instance. And while you're launching the instance, no, I don't want to walk through. I know what are the steps I'm following. So, it will give you a few things that what is the name of this EC2 instance you want.
And as I mentioned that I've selected Ubuntu and I'm using 24.04 LTS. Okay. uh for this particular SQL and what I'm doing all of these things are clear what I am using is C7XI flex large you guys can also use medium but keep this in mind that it is having enough storage it is having enough RAM so I'm using a one that is having a RAM of 4 GB which we are able to see here that is 4 GB of memory and two virtual CPU okay and the key pair logging yes I have created a new key pair that I downloaded This is nothing but ap file.
So I wanted a PM file. So give this file a name. Download this file and keep it secured. This is the same file that I used to login. So once this file is present from this directory, open the terminal and then use SSH - I then the name of the PM file then Ubuntu and and at the rate the IP address of your EC2 instance once you created that. Okay, this is how you will be able to login via PEM file. This is what I personally follow.
So this is the create pair. You will be able to download it only once. So keep this in mind. Once you create that, then you can assign some or create a new security group because we already created a security group beforehand. You can click on existing security group and select the one which you created with the inbound rules. I created the security group of logistic database server one. So I will be selecting it. So I don't have to do it again.
And then the configuration I chose here is 30 GB of storage that is none. I don't want any S3 file system. There is no system. And click on launch instance. So once you click on launch instance and then yeah the once the instance is also running. Okay. Once the instance is also running then what you will do? you will SSH into your terminal, install the Docker there and then run the Docker command that I have shared you here.
Okay, run the Docker command that I have shared you here. So, let me write the Docker instance. Okay, so EC2 instance. Once this is done, then launch instance. Launch the instance. Once the instance is launched, then SSH using PM file from your system. Then install the docker. You can Google what are the docker commands you require to install the docker. Once the docker is installed, then what you need to do follow the command to run it.
Okay. So multi-line with volume EC2 this particular command I will suggest you to run inside the container. Okay. So once you run this particular command the docker will be running and then once this is also done copy the IP address of the EC2 instance public. So once you copy the public IP address of the EC2 instance that is present here this is my instance click here. So once you copy it then you will be able to use it in the VS code.
You can use it in the IP address. You can use this to connect to the cloud legacy system just like I did. And this completes our understanding of the data. Whatever we have done so far is not the responsibility of an AI FDE. It's the responsibility of a client. But because we are here to learn something, we needed to create that legacy system. So we can interact with that and create the solution on top of that. Now that we have some ground to stand on, we can finally start the project.
We can finally discuss the solution. Let's get started. So think of it as like the brainstorming session. Okay? We are brainstorming the core concept here. So we are going to write whatever things comes in our mind. Okay. Literally. So this is not that technical part but this is about ideiation part. This is about brainstorming. So let me write what things that we have. But before that let me quickly write the topic that is solution discussion.
So first of all we are going to write what do we mean by talk to our data. Okay. So let me write what do we mean by we want to talk to our data. That is what client is going to say. Talk to our data. All right. Now uh now we are whatever we are going to do here is going to answer this talk to our data and whatever the solution we are going to present to be able to solve this. Okay, first of all let's ask first question is our data structured?
Is our data structured? We have two different parts. Data is structured yes. data is structured. No, because if the data is not structured, okay, then we need to follow some different path. But if the data is structured, then there is different path, right? So if the data is not structured, what do we mean by that? It basically means that it can be unstructured form, it can be text, it can be PDF, it can be images, etc.
But if the data is structured, if the data is structured, it means it belongs to relational database part. Right? What I'm also going to introduce you before talking about the solution discussion is that client already gave us the data. We are considering that the legacy system we have created so far is given by us from the client. So we are going to mark this as given. Now client has also given us some SOPs. Okay, standard operating procedure.
It means whenever they have a particular scenario, they read these SOPs and based on it they also take some decision. Whenever it required decision, not every query, not every analysis require decision but let's suppose this is already given to us. I have prepared one simple SOP for us to follow. Think of it as like client is giving us. So let me also paste that SOP here. So I'm going to call that as policy given by client and I'm going to keep that under policy.
So inside data we have the actual source. This is nothing but from where we have the data set from. This is not for the client. This is just for us to u track. Okay. Raw is that CSV file. Now we don't require it anymore but we are keeping it under our GitHub. But this is something that we require that is policy. So client has given us a policy in the form of MD file. I will talk about in the solution uh discussion part what else can be present there.
So we have to create a solution accordingly. But when we say client chain incident SOP this is what they given us and let's suppose this is about Southern California logistic operation. They give us few operations about standard operating procedure. This is effective date from January to 2021. and they are following this. This is only for internal operations. The version is 2.4 for whatever the documentation is because team usually keeps their documentation updated.
All right. And they have written few standard operating procedures here. First is temperature control and spoilage prevention. This is about cold chain. The name the main name of this project is cold chain logistic. Right? So what we are doing based on refrigerated fleets because when they are doing transportation we have seen a column called as temperature. Okay we are making up uh lot of things here but that is what making it is interesting because client can have lot of different details that we don't even know.
So we are making up lot of things here. Okay. So all refrigerated fillets must contain strict IoT temperature. We have those sensors that uh counts or that checks the temperature of the logistic right. So we have fresh perishables. So I temperature must remain between 0 and 4°C for this particular kind of data. Critical breaches if I temperature exceeds four then they do something they call it as breach. So what is the protocol that they follow?
The dispatcher, whoever is responsible, must immediately contact the driver to start the auxiliary cooling unit. If the ETA delay is greater than 1 hour because we also have ETA, right? If the traffic is high or the temperature is something else, ETA can increase. If the temperature is greater than 1 hour, divert the vehicle to the nearest emergency cold storage. So, this is the standard operating procedure they follow for this.
This is not present. This will only be present inside the company. These are internal docs of the company. Similarly, route congestion and diversion tactics. I've also mentioned here that if this is about port congestion heavily impacts SLA compliance. If the congestion is heavy, then the compliance the SLA will be broken. Right? So whenever we have something like this that the severity index is greater than seven. Okay, that is this particular column.
Then do not hold freight at the port. Divert all active shipments to this particular area. So again just a dummy thing but we are giving a standard operating procedure that is not related to just data but whenever we have this kind of data follow X process and the final one I have written is risk classification trigger that any shipment classified as high risk combined with a delay probability because they also have probability in their columns must be escalated to tire to logistic manager.
So let's suppose someone is asking about okay we found this data now do what they have to read the SOP they have to check with the manager what they do in this particular scenario and then do the job but when we are talking to an AI assistant or AI chatbot our system should be able to read the SOP whenever it gets this particular data and then give us the steps that now you have to contact this now you have to contact this person do you want me to write mail etc that is what we will be doing right so this is the SOP let's suppose uh they have provided So I have written two different given here.
Legacy database is already given to us. We are considering that we have already created. So this is given to us. This SOP is also given to us. All right. So considering all of this now let's proceed with the solution discussion. Again is our data structured. The answer is yes and no. If yes then structure means relational database rows and columns. No means text PDF images. So whenever we have text PDF images, how do we proceed?
Can anyone guess? You guys can guess. You guys can pause it. Write your own solution discussion here for you to follow. But uh yes, when the answer is yes, what kind of solution we should proceed with? Right? And if no, what kind of solution we should proceed with? If the data is not structured, it means we can follow a rag based approach. Yes, because we need to talk to our data. How do we get data to the chatbot using rag?
If the data is relational database, can we use rag here? When we ask rag, rag will not work because it has uh right now we have 30,000 something amount of number of rows. But we cannot get all those 30,000 rows of data into the chatbot into the AI context and then answer it. It it it is not going to work. So when we have relational database then we will ask when we have relational database then we need to proceed with something else.
Okay. So we will talk about the question but when we have relational database we need to have some kind of tools or we need to have APIs for ABS for usual analytics. Okay let me not write tools but API for usual analytics. Let's suppose analytics to check or dashboard when they use it in dashboard analytics to check the shipments amount uh per month per week per four weeks rolling four weeks etc. So they have those API analytics.
So this is the clarity that is our data structure. Yes, then what we do if no then what things we apply. So if we have things for a if we have API for that to interact with you'll be able to interact. But because we don't have API to interact with we need to find out the solution here. So let me write that and this pink highlight and others are two different parts relational database text PDF images and in our scenario we have both of them yes and no okay so we need to follow both the yes and no part because we have relational database and we also have an unstructured data that is textual data now let's talk about what kind of query they might be doing with this data Okay.
So let's write query details what kind of query they can do because right now if I talk about the API API for user analytics but are they going to ask anything if anything what kind of queries they ask this is something that we need to ask to client but for now because we don't have client I am the client and we are creating the solution we will be kind of deriving our own questions. We can ask questions based on locations.
Someone who is using the dashboard will have ability that San Francisco I want to uh use something else. I want to get the location from something else. Get the analytics for that location. So based on the location they will be quering something then based on coordinates as well because the database the shipment itself tracks the coordinate. So we require a way to be able to coordinate it. If someone has a coordinate and they want to understand within this particular coordinate how many shipments are going, which of them has high risk etc. they will be able to query uh the dashboard via the coordinate.
All right. So next thing is they can ask questions based on perishable. Okay. They can ask questions based on perishable and based on seasonal heat spikes. Okay. Or questions based on heat spikes. Okay. So, usually when we are doing shipment, right? Let me write it here. So, usually when we are doing a shipment, we sometimes get unseasonal heat spike. Uh also, and if the heat spike is expected, is our refrigerator unit in the vehicle is working correctly or not.
So, we can ask questions based on that because the temperature is there. So if not heat spike then we can just talk about temperature based question. Okay temperature or refrigeration based questions I think heat spike is going to be just too much for this whole scenario. So temperature or refrigeration based questions and also weather condition based right based on weather condition if it is raining based on the coordinate based on all this combination as well plus the combinations of a okay all of this is giving us vague idea but still you'll be able to check based on the solution whether uh if we have this API do we require new API do we ask new things from the client or do we need to create all of this on our own.
All right. So this is asking uh this is the query details also they can ask query based on SAP. So query from SOP they can ask that if they want to understand some kind of SOP that I'm new here can you tell me what are the different standard operating procedures we follow then they should be able to answer it our AI chatbot should be able to answer it. So query from SOP is also there. Okay. Because why do someone wants to uh query one SOP?
There can be a lot of scenarios. A fresh uh you can say intern that is joining the company or fresh a fresher is joining the company. They want to be able they want to understand what SOPs they follow and also the senior manager or any manager because if the SOP keeps on changing they cannot remember all the details. So there should be a way where they can ask questions or understand SOP and because we are in this AI solution we should be able to provide them this.
So what are the different things they can ask location coordinate perishable temperature weather and combinations of any of these above also the query from the SOP. All right now let's talk about the next thing. This is about query. This is about is our data structured? If structured then we can follow rag. If this one we need to find it out. But one more thing about this project that we need to worry about or any engineer needs to worry about is security.
How are we going to make sure that our agent does not insert agent does not insert or delete the data. We hear this a lot that uh agent that is deployed in this particular company uh deleted the whole database or inserted something new or interacted with the database that we don't want. How are we going to make sure that this client does not faces any such issue? Okay, how to take this? What does it take to build this with security in mind?
So let me uh write this question. What does it take to build this with security in mind with security and mind? Let me highlight that here we are talking about security. any other thing that we need to highlight all other things are good and also we need to worry about this insert delete data. So what is the solution? We have this agent okay and agent is running some unnecessary command and action will be performed by the agent.
It can be delete or insense. So how can we stop agent and we should do something before the action or while agent is taking the action? We should restrict it. So what can we do? First let's talk about what are the different kinds of uh actions agent can do insert and delete. So what we will do in between we need to find this is a question mark. What we'll be doing here we will to solve this particular thing because we have database right?
We have database. We can create multiple users. We can create a user and to this user we can restrict the access. Okay, we can restrict the access. So user will give user ID. We will have that user password. And whenever we are going to give code or access to the database to the agent we are going to use this credentials. We are not going to give the credentials of the user that has access to access to multiple things that is maybe the admin access maybe access of someone else but we are going to create a new access so user cannot interact with that.
So we are going to create multiple users based to the need and if we can create it even if agent or someone even ask to the agent can you please delete this database for me agent if it writes or executes a particular thing it will not be able to delete because that particular user does not even have access to that. So we are blocking the access from the core from the software side and even if agent wants to try that it will not happen.
So we are going to follow this particular process. Okay. Now let's talk about what are the different data sources that we have. What are the different data sources that we have here? So we have two different data sources. I've already written this before but I'm writing it now. Relational database because the more we write the more clarity we are getting. Okay. So we have relational database and then we have knowledge from the SOP. knowledge from S O P.
So two different parts we have here. All right. Now in terms of relational database, we can ask few things like are we expected to see this data grow? Yes, the answer is yes. It is a relational database the data will for sure go and to solve this what are the different things that we can do as I mentioned before API but now we will talk about the question mark part previously it was question mark so how to solve this particular data how to actually retrieve this data and pass it to AI agent now if I talk about this we can use MCP and through this MCP we can create our custom tools okay tool tools for the data.
What is the requirement for this particular step? If you want to expose an MCP to an AI agent, so it can use this tools and give us the answer and that answer we can pass it to the agent. Okay. So agent will be interacting with this MCP. If you want to have that then we need to decide the tools. Then we need to create the tools. Now when we create the tools the question arises what tools? How many tools? How many tools?
This is again a question mark for now. So previously our question mark we were at the stage of we were at this stage but now it's about how many tools because we need to ask the system we need to ask client and if the client is not technical it will take a lot of time for us to kind of derive what many tools we need to understand the data and what kind of question they ask. We have listed down the questions from our side but we really need to ask client what kind of questions they interact or they are going to interact with the system.
Okay. So this is the first solution. What is another solution that we can have here. Instead of this MCP we can also use a very popular solution that is text to SQL. companies are hiring for this particular specific uh position where ML engineers or AI engineers who know text to SQL they want to hire those individual but because we are afed to know all this. So text to SQL is nothing but if we are asking if a AI agent is asking that I want to access this data set and I want to get these many details then our AI agent is going to write the SQL to be able to interact with the database and we are going to use that dynamic SQL the SQL that AI agent has created and then run that SQL.
Okay. So we can use this particular part. But what are the cons of this? We can write that what else can go wrong in this particular path. We can say that if the query is not correct, if query is not correct, we will not get result. we will not get a result. What are the pros of this particular um direction? Not limited. If I talk about MCP and tools, we have to create 21 tools. We have to create 20 tools. So we need to decide how many tools and business tomorrow if it if they ask that we want to create five more tools because AI is not able to answer or we created or we added new data there then this particular part we need to re-engineer but in terms of text to SQL all we have to do is we have to update the description we have to update the context of this particular tool that we are creating and add that many columns there just all of that in the description and AI will be able to create the SQL query for that.
So it is not limited at all. Second is what else that is the main advantage. We don't have to write more. So it is not limited. It is not at all limited. Okay. And also whatever the cons that we able to see query is not correct, we will not get the result. This can also be solved if you create if you create a agent that uh can interact and see the previous response because if the response previously is not correct the agent can look that the whatever query I've ran is not correct I'm able to get this error agent can now create a new tool agent can now create a new SQL query and can run it again.
So agent has that ability and so whatever the cons of this particular scenario was is not cons anymore. Agent can solve it although it will require multiple turns for it. Okay. So this will come under multi-turn conversation. So we cannot just do one by one conversation. It will be multi-turn and agent should have visibility to all the previous messages. So if I compare two different paths to solve the relational database part, I personally feel like the text to SQL is here has better things to provide.
And one of the pros that we can get from here is one more faster prototype. Right? This one will take lot more time as compared to text to SQL. Yes, if you want to compare if you want to create text to SQL on a bigger scale it will take more time but for us we have one database and it is fairly simple database so agent will be able to run it we can only know by testing it I have tested it but I will be showcasing you the same thing so the solution that we are proceeding with okay this is solution discussion we can select this one okay so let me mark this as green so we know this the path that we are choosing.
Okay. So for relational database we know the answer. Now now let's talk about knowledge from SOP and how are we going to solve this. So for SOP uh when we talk about rag based solution we only need rag based solution. uh because the user can ask any queries to an agent. User is going to ask query and that query can have different meanings. We cannot just do uh fuzzy search. We cannot just do text matching. What we have to do is we really need to do semantic search.
And the best way to do semantic search is by adding that rag system. By adding the rag rag system. So the agent can now use this rag as tool. We can create rag as tool. Agent can use this particular tool pass the user query in this tool and can get the response. So if you can create the system then it can give user the context from the rag. agent then can use this context to give you the final response. All right. So if you can create some system then we'll be able to solve this knowledge from SOP problem.
Now there are multiple things that we need to consider when we think of rag. It means chunking strategy. Chunking strategy. Okay. And also when we talk about chunking strategy you also need to consider few more things like whether these SOPs are going to be updated. If updated is there going to be any modification. So modification/ uh deletion strategy. So we also need this while we are coding while we are doing something we also need to keep this in mind because these are SOPs.
We are going to fetch the SOPs from wherever is possible. we don't have to fetch it in this project but in real project we have to fetch it uh maybe once in a month or uh twice in a month whenever it got updated but whenever it is updated let's suppose whatever the agent was answering based on the SOP previously now that particular sentence itself is not present in SOP the SOP is updated so it also has to deal with modification or the deletion as well so we need to worry about what is the chunking strategy that we are You also need to worry about modification.
And one important thing in rag is I am out of space but let me write more things here that is response should contains the source information for validation. So for validation we need to have source information only. If we have source information then we can say that the agent is not hallucinating and make and like making up all this details. Whatever user is asking only if it is found then only if it is found then give me the response.
If there is no source present we can write in the prompt. If there is no source present or from the provided tool rag as a tool then please say that I'm not able to find such information. So only if we have some way to validate our source information then we can create this. So we also need to have this strategy in mind while creating the solution for knowledge from SOP. All right. And what else that we can do here is that we also have few more things.
I am literally out of space here. But let me create another arrow that is it is going to have we need to have injection injection logic for different data types or source type. Think of it as like for MD file, for PDF, for XLSX, for txt. Right now we cannot support everything. But MD file should have different injection logic because MD file has different headers etc. A PDF have something else. A txt does not contains anything.
So we need to consider it separately. Xlsx file or a CSV file should not be ingested like the MD file. So we need to have little bit of logic here as well. Okay. So this is about injection logic here. So let me write it not this color was using something else. Okay. All right. So this is what we are able to derive it from our discussion. So what we saw here we saw that let's again do a little bit of a recap that is our data structured yes no we have both of them so we need to proceed with that direction if yes uh then it is our relational database that is structured database if no text PDF images is there rag API for usual analytics business does not have or the client does not have any API for usual analytics or they are not providing us let's suppose then we need to proceed with different solution what are the queries they are expected based on location coordinate perishable temperature whether combination query or from SOP etc we can have all different queries so uh and if I talk about the security we are going to provide the security I'm not adding any guardrail because this is internal to their system guardrails is for system where uh we have customer data right now we don't have customer data we don't have to provide guardrails the security we are providing so agent cannot do insert or delete on on on the data.
So this security we are adding. So agent action itself is restricted. It will not be able to use it. Okay. And the data sources provided us the answer that we have different strategies. MCP we can use text to SQL and we are following with the text to SQL for the relational database part for the structured data part. For the nonstructure that is knowledge from SOP that we have we are thinking of a solution that is rag as a tool because from the rag we can get the context and we can pass that context to the agent.
Agent can use that context to give the user final response and based on that response also if someone is asking a question based on relational database it can check to the SOP and check if something is wrong or not. But this is about whether rag is a good tool here or not. But to implement drag we need to have chunking strategy. We need to we are also adding a validation source because if it is not coming from SOP it should not even be a part of answer and modification and deletion strategy because it is SOP it can be modified or something can be deleted.
So we also need to keep that strategy in mind. And finally right now I have only added MD support but our system should be able to read this injection or get this injection logic for all different kinds of file. We can support MD, PDF, XLSS, .txt. CSV as well. But for now, we only have u where is our data policy. So whatever policy we have for now, we can pull it here. But we don't have to pull it. We don't have any client here.
We are if we want client tomorrow if it adds PDF here. XLSS here, we can just create a new script just like this is injection script. We can create a injection script for SOP that can push the data in the vector database on the cloud and that is already updated there. All right. So we can have that strategy. So this covers our solution discussion. All right. This covers our solution discussion. We have designed the solution on an highle overview not the coding part.
This is something that engineers also do and usually uh engineers that focus too much on the text tag skips this particular part and starts implementing. Okay. So where is this notebook? So this is what we have decided so far. So what we're going to do next is that now that we have designed this solution, it's time to create uh HLD and LLD uh review uh documentation. I have already prepared that based on all of this whatever we have written.
So what HLD and LLD does? HLD is highle documentation there we just write highle information. LLD is more like technical talk where we will talk about what exact technical things that we'll be doing because until a developer has HLD and LLD they will not be able to communicate with the other teammates and can guide better. So once we have HLD and LLD this is kind of a review for the internal technical team. Once we have that documentation and that is approved, we are here.
We are going to approve it. Then we are going to create a business pitch. So business know what things we are going to create. Once this part is done, then finally we can jump to the core solution building the actual project. All right. All of this is a part of project. All of this is a part of AI formal deployed engineer that they have to do which is separate or different from other engineering. uh you can say flows.
All right. And that's all. Uh now we'll be jumping to system architecture. Now that we are done discussing the solution, it's time to convert that discussion into the actual technical document. It means using this document, the team is going to understand what exact things we are going to build in this particular project or the project. And once we have this document we the team leader or the AI forward deployed engineer can assign this task and the team will have a better visibility on what actually we are building before even proceeding with the actual solution before even writing a single line of code.
This particular documentation helps the team as well as the business not particularly business but if business has some technical person this can this document can help them but it is usually for technical team to understand yes this is on paper and all of those things is manageable all of those things can be coded can be created so finally let's create this documentation proceed it from there and build a final project so let me share my screen where I have already created this technical design documentation TDD for our project where it has high level document as well as lowlevel document for our project.
So this is technical design document TDD cold chain logistic AI assistant and the author is forward deployed engineer AI FDE. The status is this is architecture review. This is not the final thing as I mentioned. This is still a documentation but it a AFD or a AI engineer or a software engineer usually creates these documents once they are done with the discussion that we have already uh did here in this particular part and from this discussion because lot of things are not technical here we mentioned rag here we mentioned database but we not but we have not mentioned the exact things lot of the exact things we try to build it in this particular stage in this particular document so we have also mentioned the date that this has been discussed in this particular month and these are the table of contents for us to follow.
Okay. So it will have system architecture and system architecture only has two different components which is a part of HLD. Then we have agent orchestration etc that is a part of LLD. So all these are part of LLD. Then we also have deployment strategy. We have already deployed one thing but it's more about what we will do with the actual solution because database and legacy system is something that already exist although we have created here in this project but now it is kind of that something is given from the client right so system architecture HLD now let's proceed with that so this is the cold chain logistic AI assistant is decoupled three tire enterprise platform designed to enable control tower operators to query fleet telemetry this is data a telemetry data that we have environmental conditions and compliance SOP via a natural language interface.
So this is just a fancy term to say that cold chain logistic is a AI chatbot for that operates on query telemetry data environmental conditions and compliance SOP. These are the three different tools that we are going to use it in this particular AI assistant. So as I mentioned that we have three different tools here that is the same present in this particular architecture overview. So let me zoom in. The system relies on a strictly stateless presentation layer a state machine driven orchestration engine and a hardened data access layer that is a security part that we have talked about.
So user the operator is going to operate with streamlid and the streamllet is going to use langraph agent orchestrator. It is going to use the orchestrator and orchestrator is nothing but that agent state the full state that we are creating the reasoner etc. I will be talking about that but that is nothing but a lang graph component and that lang graph will have access to tools the tools that we are going to create SQL server that is nothing but query fleet telemetry data our text to SQL part vector database is nothing but our compliance SOP and rest weather API that is our environment conditions why I have added this new tools which we have not talked about in the discussion well I just want to give you clarity on in real world.
Whenever we have this legacy system, if someone wants to understand the temperature outside there, right, not just a data based on the sensor because sensor will be giving you data information based on the sensors inside the vehicle, right? So, we need to understand that if there is uh if it is too windy outside or it is too hot outside, we can check based on can you please scan all of this uh latitude and longitude and tell me which one is having that concern.
So we need to know the real information there, the real weather there. So I've added this new thing. This is a external tool that we'll be using, external API, but this is just a good add-on to uh increase the complexity of this project. Okay. So this is a highle architecture flow of this project. What are the different components here? We have presentation uh tire that is nothing but our streamlid application itself.
And using streamlid we are going to manage the UI. We are going to manage the session state. All every reasoning is offloaded to the orchestration layer. So presentation layer is just the UI which manages the state as well. Now coming to the second layer that is orchestration engine. This is going to be a langraph react agent. As I mentioned that our agent will be able to uh get the answer from the tool and based on the reasoning it can call itself again.
So this is called as react agent. This component is the brain. It manages the execution life cycle including intent deconstruction. We are going to find out what is the intent of it. Then multi-step tool selection. Agent can use all three tools. One tool based on the problem or based on the query user is asking and synthesize of tool outputs into human readable responses. That's why we are using AI chatbot. Right. Final thing is data and integration layers.
So we have decoupled backend system into multiple systems and the system strictly isolates the LLM agent from the raw underlying data by utilizing database views. So instead of creating uh yes we are going to create users but that user is not going to have access to the main table it is going to have access to database views and only read only role rappers. So that role we are going to create is only going to be read only.
Okay. Then we have agent orchestration. Now we are finally going inside the LLD because the streamllet part we don't have any complexity here but the complexity arises in this particular part that is orchestration and data integration layer. The orchestration tire ensures deterministic behavior with a probabilistic framework by utilizing a state machine architecture. Okay. So uh we are going to manage the state and execution loop.
So we are going to have injection and passing. Okay. We are going to ingest the SOP. We are going to parse the agent receives NLP. So we are going to parse it and then specify the specific domain whether it belongs to telemetry this query response SOP or whether. So based on it now it will go to the reasoning loop. The reasoning will check evaluate the current graph state. If the intent is ambiguous it requires clarification.
It will ask user that can you please give me clarity on what you are asking. If it is extinctionable and one of the tools can be called it dispatches to the optimal tool as I mentioned. Okay. So this is more about the langraph. If you are not familiar with the langraph you can skip this particular part. But AI FD engineer as I mentioned till this part there is no complexity on what exactly we are using. This is more like a formal discussion what we are building.
But once this is done then a AI FD engineer will be doing some homework on what exactly and how exactly we can do it. And let's suppose lang graph is the tool is a perfect tool that we can use then we have included all that in this particular LLD. Then the tool execution the agent is going to invoke the tool isolated worker functions ensuring no side effect occurs outside the tool specific code. So tool when we invoke a tool is going to give you the output that output will be sending we are sending it to go it to a reasoner and agent is going to take the reasoning from that.
Okay, the uh step from that then the synthesis once we have the tool output then as I mentioned that it will go to the reasoning is synthesized by the LLM into a cited context aware answer that from where we are getting this answer so we are going to have that I've created this sequence flow this is more like a technical one where the user is going to ask agent is shipment A42 on track we don't have something like A42 this is just a example because only after creating the core functionalities we will know what kind of things it can answer but in the database we don't have anything like A42 so is shipment A42 on track it is going to ask to agent then agent based on this one is going to select which tool we can utilize it for if not it will ask user or it will tell user that I don't have this information once it gets this query then it will say this first let's uh check if the shipment ID A42 is present let's retrieve on the data so it will create the SQL query pass to the tool.
Tool is going to run the SQL query agent created and then uh it is going to get the answer. Okay. So here because uh this is the diagram generated. So from the query database it is going to get some answer and it can also check it can also reason that okay I have the information about shipment ID. Now I also want to find out what is the latitude and what is the temperature etc. I also want to get that. Then it will go to the external source API get the result back and then it will say yes it is on track but because the temperature itself is too high maybe uh temperature is feeling unusual lets me check the SOP that I have the information of the shipment I have the temperature do I need to do something on that so now it will go on the fourth step that is to search the compliance SOP for the temperature range because we got this temperature range if the temperature is saying that yes it is in limit SOP saying that temperature is within range then it is going to return the final answer of all the tools etc and going to reason that yes shipment A42 is within range minus 20.5° and also it is on track etc.
So finally user will get the answer. This is the whole thing that we are going to create from the lang graph. All right. And this is the LLD for that state management and execution loop. If it feels too complex as I mentioned the complexity of the project is high. Don't worry about it. Just understand from this one that the agent is going to ask get all the details and get back you the answer. Okay. Then comes the resiliency and fallbacks.
So graceful degradation. If the external API calls weather fail, the agent will report and will not take a hallucinated reasoning that uh it will it is going to say that the report is missing rather than hallucinating the status that everything looks okay. It will say I I whatever reasoning I'm doing uh it is without this weather that we have selfcorrection. This is what I mentioned before that if a SQL query returns an error or because this is one of the uh cons of using this particular approach where is that that query is not correct we will not get the result.
Then because it is a react agent it can do the selfcorrection. If a SQL query returns an error the agent logs the error and attempts a simplified query iteration before escalating it that the database is not accessible or something is wrong with the database. It will try multiple times and we'll do the selfcorrection. So this is about resilency and fallback and once this is done then it's more about data and security model LLT security is implemented at physical database layer.
So we are not creating anything that can be broken. It is present in the database layer in the physical database layer to mitigate prompt injection risks. So we will not be able to do prompt injection. anyone who is doing prompt injection with the LLM because they don't even have access they will not be able to do anything okay the agent itself does not have access the user of the database does not have access so we are going to uh add that security telemetry security I will be adding that uh in a bit not in this particular section but I will be adding that and how to add that the query we will be writing that security principle by exposing only the flat view so we are going to create a view we're going to create a user that has only read access to this particular view.
We physically prevent the agent from executing drop, update or delete commands. Okay, physically here does not means that we are going to have a physical robot but it means that from the database layer itself. Okay, then as I mentioned role based access control so we are going to create the user user FDR. Okay. And we will only be giving this access to uh view okay or to only read. So account this account is going to lack permissions to write modify or access tables outside the specified view.
So we are going to create that and immutable auditing. So for this particular project what we are also doing because lot of the times client does not want say data to send to some other resources or or some other u softwares. Okay. So what we are going to do we are going to audit or we are going to log each and every step of the agent in our database itself. So we are going to create that new table in the database that is agent audit log.
Okay. It is committed to agent audit log and we are going to log each and every entry. So you don't have to create account to the lang views or send your data outside your system. Your logs and the system is also present in your own database. So nothing gets out. All right, this is about log entry. We are going to implement that as well instead of using lang fuse. Think of it as like the requirement from the client. If your client is okay with lang fuse, just implement lang fuse.
It is easier than creating all this new schema etc. Okay. Then the deployment strategy what we're going to do we are going to use infrastructure as code approach and this one is particularly is going to give you a better idea. As you already know that we have created a database server MSSQL server inside a docker. This is nothing but a private only. I will be able to access via SSH only my IP address. Also if you see one more thing that I was telling you that I will be teaching you later that is if you go to this particular EC2 instance.
Let's see how many instances that we have. And if you go to coal storage logistic SQL, you will see that this one has access. Let's look at give me a second. I just want to look at not the sub subnet but okay not this one as well. Okay. So let me show you that part. So what we were using it before that is security groups. So for this particular EC2 instance I have added security group right I've told you what was the security group.
So while you create this database you can uh in this EC2 instance you can add the security group and if I search for security group okay if I search for security group where is that not here. Yeah if I search for the security group for the database server that I created. If you see here, let me go back and actions view inbound rules. Okay, so I added another security group and I'm giving it access to access my EC2. So other than my IP address, only this security group has access to access my database or access my EC2 instance.
And what is that security group? That security group is nothing but we are going to create a EC2 instance and to that EC2 instance we are going to add this security group that I created. So if you see I have multiple security group here. This one is for streamllet. So this security group is nothing but this is streamllet security group and here I have added let me show you that. Yeah, here I have added SSH from my system and then the custom TCP from 8501 which is nothing but this is where the streamllet runs.
Okay, this is the port where streamllet runs. So in this EC I'm going to create a new EC2 instance. Assign this security group to that EC2 instance. And because this is a security group named this I added this security group in my database. So only this can access my database. So this is internal within the AWS and this is how it looks like. If it looks confusing, you will be able to understand it from here. So I created the database SQL server.
I can only access this from my IP address. Other than my IP address, only this particular uh EC2 instance can access my database because this one already has a security group assigned to it. So this security group can access this security group. Okay, I have added that configuration there. This is virtual private cloud. This is public subnet because public will be able to access it. This is what I have mentioned here.
My SQL custom this security group others I can add another rule I have not added. I will tell you. Okay, let's add another rule. Custom TCP we want people to be able to access this anywhere. Okay, if I add this then people anywhere will be able to use the port. Sorry, I think this is different one. We need to add this to the app server. Edit inbound rules. You can see yes. So 8501 is accessible from anywhere around the world.
This one is already accessible to the public but only this server can use my database. Okay. And that is what is present here. So this is internet from the internet will go to the internet gateway of the Amazon from the Amazon they will land to streamllet that is what user will see this streaml can also see the security group or the components or the systems present in other security group that is nothing but my myql server.
So this streaml has access to this. It can uh query this database. The tools will have access to call this particular database and get the answer and user will be able to see the answer in the streamllet. So everything is protected. The app node is nothing but AWS EC2. This is AWS EC2 AWS EC2. But the orange one is app node stream EC2. Data node is nothing but this MySQL EC2. Okay. Green one is public access. This is 8501.
Public can access this using 8501. And this one is private access. This whole thing is private access only internally from this. So this is more of the infrastructure topology that we have created and we have also written that network parameter security ingress public only port 8501 is accessible from the AWS EC2 this streamllet and the internal communication port 1433 SQL is
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.