
Advancing AI Datacenter Cooling: Intel's Intelligent Rack Liquid Cooling Control Center Architecture transcript
Open Compute Project · @OpencomputeOrg
Words
1,667
Runtime
15:34
Speaking pace
107wpm
Reading time
7min
107 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Okay. Uh so uh good morning everyone. Uh I'm Jay Wu uh from Intel. So and uh it's great to be here again uh at OCP Epic Summit this year. Uh today I will share our work on uh intelligent rack scale liquid cooling for high power AI infrastructure. Uh we will start uh
54 words, the words spoken in the first 30 seconds at 107 words per minute.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 92 |
| Average words per sentence | 18.1 |
| Longest sentence | 66 words |
| Questions asked | 3 |
| Sentences containing a number | 1 |
Most used terms
- uh201
- control21
- liquid17
- cooling16
- components15
- rack14
- level13
- architecture11
- different9
- okay8
- system8
- intelligent7
Filler phrases
207 in total: uh 201 · like 2 · um 2 · actually 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Transcript
Okay. Uh so uh good morning everyone. Uh I'm Jay Wu uh from Intel. So and uh it's great to be here again uh at OCP Epic Summit this year. Uh today I will share our work on uh intelligent rack scale liquid cooling for high power AI infrastructure. Uh we will start uh with the challenges we see as AI rack power continues to increase and then uh walk through our control architecture implementation and the key learnings from our proof of concept.
The main message uh is how we can move from individual liquid cooling components toward a unified intelligent rectable cooling system. So uh in addition uh to the presentation today, we also have a live demo rack in the showcase area. Uh my colleague Donghan uh he will demonstrate how the solution works uh in practice including the monitoring and intelligent control functions. So I will welcome you to visit our demo on the four uh fourth floor after the section.
[clears throat] Okay. So, uh let me first give you a quick overview of today's discussion. Uh we will start uh with a radical cooling challenge and overall architecture. Uh then we will uh look at how we use mobile as an open uh communication interface to connect different cooling and infrastructure components. After that uh we will walk through the hardware and software management architecture and how we deploy at the rack level.
Uh finally I will share how the system handle real world liquid cooling events including uh leakage and uh thermal anomalies. Okay. So let's get start. Uh as AI workloads uh continues to drive higher compute density and power consumption, the radical cooling uh is becoming increasing uh important for nextg AI rack. But cooling the rack is not just about putting uh uh co-play on the CPU or GPU together. uh as we add more liquid cooling uh components such as the memory switch and other infrastructure the number of the cooling components and liquid loops increase significantly.
So this creates a new challenge. How do we monitor and control all of these components as one integrated system? So today uh many deployments still rely on different vendor tools and individual control interfaces. Leak handling can also be manual and uh coordination between the CDUP PDU valves and sensors is often limited. So our focus is to move from component level liquid cooling to intelligent rack level control with unified temperature event and automate an automatic uh response.
This is the uh motivation behind the radical coding control center that we've uh developed. So this page uh shows our overall architecture. Uh we bring together uh the major cooling and infrastructure components including the CDU, PDU, uh flow and pressure sensors and thermal sensors uh and also energy control valves and leak detection uh components. Instead of managing each components independently, we bring the termetry into a common control uh framework.
On top of that, uh we implement rag and no level intelligent control uh for event detection and response. This key uh the key enabler here is the use of open communication protocol including uh mobile RTU and uh TC P TCP IP interface uh which give us a particle way to integrate components from different ecosystem partners. Our goal is not uh simply uh to build another monitoring dashboard. Uh the goal is to create a closed loop control architecture uh that can detect an event, understand where it occurs and take the appropriate actions.
So one important questions how do we connect all these different devices inside of rack. Uh we select mobiles as one of the key communication interface because it is an open and widely adopt uh in industry. uh for uh fuel level devices such as sensors, VAS, VPS uh we can use mobile RTU over uh ICE 485 uh which provide a reliable communication and good noise uh immunity for higher level integration with servers or N KN or RE uh management.
Uh we can use mobile TCP IP over Ethernet. This provide a relatively simple and open mechanism to bring different cooling components into a common termetry and control framework. So now uh let's uh look at the hardware implementation. So at a red level uh we have the CDU, PDU, MC uh and also power infrastructure and liquid cool component and and and liquid cool computer nodes. Uh on the cooling side we monitor parameters such as uh fluoray uh uh liquid pressure and also uh liquid temperature. uh while uh the sensors and energy control valves provide a protection and control uh capabilities.
Uh this important point is uh that these components uh don't have a don't have to come from a single vendor. By using common interface we can integrate different ecosystem components uh from different uh partners into the same rect control architecture. This give us the visibility uh to evaluate different supplies while maintaining a consistent management and control framework. Now we have uh components that uh have uh temperature enabled.
Uh moving from the hardware to the software side. Uh we design the system with a uh layered architecture. At the button uh at the button is device communication uh followed by data acquisition and monitoring. Uh at the top uh we provide uh visualization and control interface. Uh this uh separation is important because it allows us to add or replace individual devices uh with without redesigning the entire system. Uh it also makes uh debugging easier uh because uh we can isolate an issue at a component level or no level or uh rag level.
This architecture gives us a path to scale the solution as the number of liquid core components and loop loops uh increases in the uh rack and also the data centers. Okay. So, uh now uh we've talked about the hardware and software there. Uh lastly, uh we are talking about uh how we handle uh the thermal event uh inside the rack. uh wireless system uh continuously collect the temperature data from the CDU PDU thermal and pressure sensors and also the leak sensors. uh when an abnormal condition is detected uh the control logic uh determines the appropriate response.
Um for localized leak for for example uh the system can isolate and affect liquid loop or note and coordinate the corresponding power response. So with the hardware we uh enable we can isol isolate the liquid loop for a single node while a uh leakage detected in that specific uh several nodes uh without uh shutting down uh the entire uh the CDU and also the entire uh decoding rack. So this allows us to minimize uh the impact of a failure while maintaining uh operation uh of the uh remaining systems uh wherever possible and I think this is an important step toward making liquid coding more silence and easier uh to operate at scale.
And so as I mentioned in earlier uh so we have a uh uh functional rag at the fourth fourth floor uh at Intel uh booth. So uh welcome uh yeah I I invite you to visit the Intel booth to uh see the live demo there after this session. So uh finally uh this is a uh the user interface uh for the control center. So we create uh this dashboard provides a consolidated views uh of the rest cooling thermal pressure and flow rate and power and also the uh leakage status. uh instead of uh engineer checking multiple vendor tools uh we can bring uh the key information uh into one in uh environment.
So actually uh we uh use the AI tool uh like copiler to build build this uh dashboard uh which is easier for for uh thermal engineer near right now we don't uh need to equip a uh uh skillful software technology but we still can de uh the AI tool to develop this uh dashboard that help us to easier deploy uh in the rack. So uh this also gives engineer a common platform uh for real-time monitoring and event analysis and debugging.
So uh the long-term goal uh is to move beyond uh v visualization uh to more intelligent uh automatic automatic rag level control. Okay, so call to action. Uh to summarize uh we have demonstrated a red devil liquid core control architecture that integrates uh CDU uh liquid floor and uh pressure and uh thermal energy valve and also uh le temperature into a unified framework. Um the multi-layer control architecture enables us to detect and localize abnormal events at both the rack and no level and and then uh take automatic uh corrective action.
Uh just as uh importantly uh the layer software architecture and open communication interface provide a path uh to integrate components from different uh ecosystem partners. Uh our next step is to scale this approach from single rack validation to higher density and uh part level deployments uh while continuing to work with uh CD vendors, sensor and infrastructure vendors uh to improve uh interoperability. So uh automatically our goal is to make liquid cooling not only uh more thermally efficient but also more intelligent uh resilient and easier to deploy at scale.
So uh that's my presentation for today. Thank you for listening. uh this work reflects collaboration across many teams and many partners uh in the industry and we hope it helps you uh in your design. So I happy to take questions or discuss offline after the section. Thank you. Okay, there's a question. Yeah, I have a question since this is the open platform, right? Any anybody have the the protocol we can link to your platform?
Yeah, like we are sensor maker. We can connect to your system. Yes. Thank you. Okay. If there any question you can go to the boots the on fourth floor and more discussion can be happen there. Okay. Thank you Jay.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Use this transcript
Three free tools that work on the material around a video like this one. No signup, no login.
Hook Analyzer
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Policy Pre-Flight
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Channel Skill Generator
Read this channel's public videos and transcripts, and download a writing brief for it.