Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Gonka // Decentralized AI · @GonkaChannel
Words
10,454
Runtime
1:11:57
Speaking pace
145wpm
Reading time
44min
145 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Okay, I think we're going to get started now. Uh thank you guys all for for coming. Uh so Ganka reached out to us maybe like one or two months ago and we thought that uh they had a very exciting you know workshop to present to us. Um so just really quickly to introduce Ganka um Ganka's team uh Gleb and Anastasia will discuss Ganka as a decentralized uh AI compute protocol
73 words, the words spoken in the first 30 seconds at 145 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 426 |
| Average words per sentence | 24.5 |
| Longest sentence | 195 words |
| Questions asked | 43 |
| Sentences containing a number | 19 |
Most used terms
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Okay, I think we're going to get started now. Uh thank you guys all for for coming. Uh so Ganka reached out to us maybe like one or two months ago and we thought that uh they had a very exciting you know workshop to present to us. Um so just really quickly to introduce Ganka um Ganka's team uh Gleb and Anastasia will discuss Ganka as a decentralized uh AI compute protocol why it exists what problems it addresses what constitutes constitutes the centralized inference and the importance of comput uh verification and I think that they can go into more detail about all of that so I'll pass it off to you guys. >> Hi thanks for the introduction.
Uh so yeah as just mentioned we are part of the team that launched this protocol a bit more than a year ago and today we'd like to have uh maybe a broader discussion not necessarily completely related to the protocol but about decentralization in AI in general and uh why would you even consider uh the concept like this why it can be interesting and viable what kind of problems it solves but also what kind of open engineering and research questions still uh are unanswered there are currently being solved.
Um and maybe some of them would be also interesting for some of you to try and participate in to do some research on. So uh to begin yeah why would you even consider decentralization? Well, um I think being in Berkeley in 2026 is kind of like probably you don't want to hear another speech about uh how AI is becoming a more and more important part of our lives. But it does affect our lives indeed. It does affect the speed of progress.
It does affect uh who can do research and who can generate uh some new results with this research. So I'd say that access to frontier models is pretty important. Now uh if we're thinking about who controls the models, yeah, if we're thinking about closed source models, surely yeah, who can uh who controls the model controls access to it and also controls the output. So what can be produced by this model? Is it anyhow censored?
Is there any advertisement ejected into it or maybe is it adjusted to serve some propaganda? Now what would be an alternative? Uh one like obvious thing that comes up that's been discussed a lot lately is open weight models. Now uh open weight models they do solve some of these issues but they rather shift control from those who own the models to those who own the compute because uh if you don't have access to compute to run your model it doesn't matter that the weights are open and if that open weight model is being served on some hardware that you don't have access to you are just uh getting it from some provider who is claiming to serve that open source model.
You actually have no control uh over understanding whether the model was modified or not. Uh whether the outputs that you're getting uh were not any anyhow adjusted and whether you didn't get for example a quantized version of it that's produces much lower quality results but the provider would never tell you about it. So compute providers they still have those same uh control options yeah over access and over what kind of uh modifications can be done to the model that they serving you.
So I think >> yeah uh so we discussed this shift from uh model to from home model to who own hardware but let's now talk who actually where actually this hardware can be. So we all know that there are many data centers built uh last years and they're pretty expensive. They're pretty heavy. But still, let's look at our life and we have compute resources at our homes, our laptops, our smartphones. And if just uh compute all this together and sum all this together, it's huge amount of compute resources.
It's huge amount of CPU, RAMs, even VRAM in many smartphone. And our smartphones are pretty powerful. And let's just come take a look on simple example. So Bitcoin uh Bitcoin miners together operates more than 26 gawatt of data centers. It's like so new data center for AI Colossus 2 is only about 1 gawatt. So the amount of uh compute capacity collected together in such in bitcoin is just huge. is more than openi has.
It's more than meta has and it's really can it's it's really beacon can be used for something useful and what is also cool about this that uh it's almost impossible to somehow limit access to such compute capacity for example you can uh say openai not to serve models to some people and they will be will have to not serve. But when it's distributed network, it's just thousands of different companies. It's different places.
Some companies are fully anonymously anonymous and no one even know who is this and it's impossible to block. You block one small data center but Bitcoin will not change from this. And if you would, so the natural question here what is to aggregate such capacity such compute capacity and actually use it to serve models to train models and actually do something for AI on such compute capacity. Yeah. Uh so uh yeah I think >> uh when you talk about distributed compute um since you use Bitcoin as an example is this like different people who each might have like a personal GPU and then running like some program across all of these uh GPUs that different individuals might own. >> Uh for example in case of Bitcoin yeah you can have GPU at your home and it will be it will have small probability of mind block and one in year you will get this reward.
So for sure when we are talking aboutization of AI it's a little bit different requirements because you actually need to have something to something to be computed pretty fast. Yeah we will go deeper in this part a little bit later. Yeah thank you. So yeah again we have several problems with access to AI uh access to models and hardware and we have pretty good opportunity of huge amount of hardware everywhere. Let's think what we actually would want to do with all this.
So yeah uh we wanted to build something which can be used for hosting and training best models on such aggregated hardware. So what would be natural expectation from such uh from from such platform. So first it must serve best opensource models with real native uh performance. So you should not wait for your response for hours because it will be not really useful. It should be fully self-sufficient. So it should not depend on the fact if I don't know some companies will stop to produce open source model.
Such system should still keep going, keep work and produce new models by itself. And also uh one more requirements is that we actually want that no one can limit access to such system and also no one can actually change the output of models served on such platform. So it will be uh quite predictable. So if it claims that it serve model A with quantization B it will just work this way and no unpredictable results no worse quality in one week how entropic does yeah and we have actually conditions and with which we work when we consider building something in such limitation so such hardware will be distributed around the world it's different servers because we'll have really poor connection over internet instead of like envink and all this stuff and also this hardware will be heterogeneous so it will be cuda IMD some custom I6 and we actually don't can predict which hardware it will be uh yeah and also to gather together really a lot of hardware uh we must allow everyone to join such network and to contribute it and they contribute their hardware.
So it's actually should be permissionless but also when we allow anyone to join uh it means that not everyone will be honest from people who who try to join hardware. So such system also must be trustless. It must not just assume that uh hardware provider will be honest. Yeah. And that's kind of the goal how we want to build in which limits we want to build the system. >> Yeah. So it it might sound tricky. Yeah. On the one hand we want reproducibility.
We want verifiability. We want to ensure that nobody altered the model. and yet in a system where anybody can kind of try to do whatever they want. So how how do you solve this problem? How do you understand what you can trust when you cannot trust anybody? So it's actually a well-known common problem from game theory. Yeah. About Byzantine generals basically like the formulation of it is about generals that are deciding whether to storm or not to storm the city.
And uh some of them, a majority of them are loyal, but a small group of them are uh traitors who are trying to uh get them to do their own um their own plan that will get them all in trouble. So anyway, how would you solve such a problem? Yeah, all all of the strategies rely on the assumption that a majority of those generals would be honest that they would be loyal. Similarly, all decentralized networks, yeah, they they are based uh all decentralized systems, they are based on this honest majority assumption and the strict requirement of it is uh the design fall tolerance assumption saying that uh we're assuming that 2/3 of the waters are honest.
Why is it the strict requirement? because it requires everybody for any decision uh for at least twothirds of the participants of such network to agree on the same result. So if more than one/ird disagrees yeah you cannot reach consensus. So uh that's like a common thing to apply. But now the question is so how do you count the votes? What is one vote? Because in case of generals from the problem surely one person casts one vote.
But if we're talking about system where you don't understand who your part who the participants are. Yeah. What would actually be accounted? How is your vote weighted? Uh and then there are different approaches to this. Yeah. One is you can just freeze money because the assumption is that if there is a lot of it then a small malicious party will not hijack the consensus would not take over more than uh onethird of it.
Now, uh if um the alternative way yeah kind of mentioned uh in bitcoin earlier would be proof of work when the weights are proportional to the computational resources that are benchmarked by some type of work. So that's uh that's the way aligns the weights and the rewards and everything in the system with the computational capacity. But majority of such system yet they have uh one big issue that the work that they are producing to benchmark the capacity is useless and is uh just burning electricity basically.
So uh when launching Gonka the uh mechanism that we adopted the proof of compute is a version of proof of work mechanism where uh the weights are proportional to the uh computational capacity that's been benchmarked by the model inference. And the key point here why it's not actually producing uh any significant amount of useless work is that this benchmarking phase is done really short that's used to like really understand what the capacity is and for the rest of the time uh the system serves inference of the open source models that participants agree upon.
Uh so uh yeah a quick recap because we we did a lot of like overarching uh big questions and what's fixed so far. Now we just discussed yeah that we uh are trying to solve a problem for the system that is distributed uh permissionless trustless we're operating under the BFT assumption and the weights in such system would be proportional to compute capacity that is being benchmarked uh on inference. Now the votes yeah that come from this weights they are used to decide on which models can be served which models can be trained uh and then the workloads and any rewards are distributed proportionally to the uh weights that are gotten from such benchmarks.
So this is kind of the the preset. Now what are actually the interesting open problems that come from from that point? >> Yeah. Uh let's go a bit deeper in proof of compute itself and security model model we have here. So actually which problem proof compute solve is how to assign weight to different participant in such system proportionally to compute capacity they have but also we are not really interesting to some abstract compute capacity like oh let's measure how many hashes it can compute u so it's not that it's only compute capacity which is specific to inference of large language models.
So it's probably we all know it's most just pure compute but it's also compute of some uh really low precision numbers and also it's a lot about access to memory because large language models require really fast read data from VM. So yeah uh how we are solving this uh once in epac we have that this p uh which has two phases first is generating fast all host simultaneously and it's it's really important that all host uh hosts do this simultaneously uh they generate some artifacts uh these artifacts actually based on forward pass with randomly seeded input of some large language model for example like Kim K3 you get this model you replace input with some seeded uh input random input and make inference and save certain artifact save it on disk.
So during this couple minutes you produce as much as you can this artifacts and everyone else also produce. Then this minute finished and every host started starting to validate each other. So they sample some random uh subset of such artifacts from another hosts and repeat and validate to check if honest or dishonest. Then they directly vote and decide on the weight of this host for the next epoch. So this like 3 minutes procedure finished and everyone agreed on wait for next 24 hours.
Yes. >> Uh yeah. What is the problem of this part? >> Yeah. >> Um what is an artifact? >> Artifact uh in this case uh yeah so they they have nons nons is like certain int and they it's like used to seed input for the model. then they make compute and during the compute they there are several random layers added to the LM and then uh like additional random linear layer at output then some random coordinate sampled from this uh layer and it's just like artifact itself is uh 12 dimension vector uh which together with input nons like allow to uh verify all this end to end phase is all the procedure.
Yeah. Uh yeah, again which problems we have here? It's actually good that we discuss this question. It will be simpler to discuss problem. Uh so computation on GPUs are nondeterministic and that's like every time and also when you have hogenous hardware it's becoming much worse. Uh so you cannot just repeat the same computation. it will be slightly different uh even without real sampling tokens just like matt mu it already will be nondeterministic uh so all verifications must be based on statistic so you can compare artifact by metric and then you should somehow decide correct this artifact or not >> but for validation when you say that each host generates random temples for each other artifact does this mean that they is generate the seeds that get used to generate like the random layers for like a different pose or >> yes uh when uh generation start seat becoming so what is used for seat actually it's the hash of block which triggers uh P procedure it's almost uh random then uh to produce final seat every uh host adds to this uh hash their public key and also So some int like nons this nons and then uh during the validation they know all seeds of each other and they just sample like oh I want to verify like this 100 from this 10,000 then they request full artifacts from the uh host and just repeat this uh computation to check if matches or not. >> So the model is randomized per each PC procedure but then each participant has their own seed.
So the computations that they are doing are also participant dependent so that everybody has unique set of nons. >> Yeah. >> Uh yeah and by the time that validation happens all the seeds are revealed so everybody can reproduce. >> Next yeah again about problems. So nondeterministic computer already mentioned also models are really optimized for different hardware and uh if you use different models for such test it will be different weights.
It's pretty complicated which model to choose because it will be biased to some certain type of hardware for example latest GPUs or older GPUs. Yeah that's kind of makes our life uh harder. Uh also artifacts are heavy so you cannot just exchange all artifacts. It will be like gigabytes of data but also you must commit to some artifacts. So you cannot compute them only when someone requested them. So it's uh yeah >> and to be clear about what art is I think you mentioned that you add like random layers at different parts of the model and then you also add I think I think like a random linear head at the at the end and then you sample one of the logic from that random head.
Is the artifact like the activations from the random layers >> or is it just like the random logic that you like pick out at the end? It's so you have input uh random and you have also with the same seats all random layers and then you make forward pass of this model with random layers on this input and yes you have like output vector uh activation uh >> oh are are all of the layers randomized or or how would the layer randomized like do they just get random weights instead of like the pre-train weights or >> uh It's additional uh really small lightweight layers uh which are not used in model outside of PC.
It's so why this layer exist because we wanted to avoid situation when someone wanted to uh distill such model and actually replace real model with small models. So just to as additional layer of security we uh seed intermediate layers also. >> Yeah. So you don't you don't know model in advance to optimize for it. >> That makes sense. >> Yeah, >> I see. >> Uh >> um but for like the random layers that you add in, do you then take like the activation outputs from those random layers and use those as the artifact you use? >> No, no, they just used to during the compute. >> Oh, I see. >> Yes.
And only final output uh vectors after all layers, both random and not random. >> Oh, I see. >> Yeah, >> layers are the same for everybody. So >> um >> I see that makes sense. >> Yeah. >> Like each random uh each of the random layers is inmed by the seat, right? >> Uh yeah. And one more problem that uh we should account for uh separately for different phase of inference for like real users. It's like LMS have prefill phase, decode phase and also it's quite heavy use cav cache and benchmark must account in some way for all of this.
Uh yeah so once again I already about security model I already mentioned it. So host claim their weight proves this weight to the network and this takes like minutes. Then protocol assign task to each of these host according to their weight. So if one host have 10 times bigger weight than another, it will have 10 times more task. So they assign task protocol assign task to different host and host must successfully execute this task. uh this task are just user inference just something some some some part of training actually it's real task it's not artificial task it's just someone go to gonka and say yeah I want my agent to work on some model and this is actually a task uh yeah host must execute all this task to be rewarded and they rewarded on the end of epoch only when all work is done and no like invalid work is being validated in some way and it must be done honestly.
Uh yeah uh the second big part about uh what I want to say to to talk is like inference itself. So again always whole epoch uh hosts are serving inference and they serving like millions of requests per day. uh but how to so the problem of validation of this inference how to verify that specific uh sequence of token uh was generated for some prompt uh with some parameter but by this specific model. So how to understand if host didn't use smaller model, didn't use quantized model, didn't use something else.
Uh it's actually pretty complicated uh to do it like with good quality. Uh second issue, second problem is reliability. how to actually verify uh that host with compute capacity w use all this compute capacity for inference and didn't for example uh serve inference for some another platform or sell it on open router or something like that and another problem it's more about engineering problem uh where to record this all again we have like network which doesn't have single source of true doesn't have the controlling part And but you need to record every payment for every inference.
And also when some decision must have to be made for example to invalidate something. Uh as we discussed uh many host must vote to agree on some decision like this majority super majority which Anastasia uh talked about. Uh yeah the question is how to record the soul? how to where to store the soul. Uh yeah uh let's go deeper in validation. So again we wanted to verify that uh the model which supposed to be used was used for this request.
Uh again actually we have similar problem with PC GPU uh computation is not deterministic. So you cannot just send same prompt to another uh server and check if sequence of token matches. It's almost impossible and happens really not often. Uh also the recomputee itself is pretty expensive. So you just need to do the work twice and also if you want to invalidate something you must confirm it by majority. So you will have to repeat not two times but like 100 times which makes overhead on such system so heavy that it become just useless.
Uh yeah and what can be solution for this? Uh so the system which allow uh which relies on repeat inference to verify there's quite well-known uh paper top lock uh which does it and requires it requires again repeat this request but it works and there is also some more research approaches like the KML like zero proof systems to for machine learning they provide some proof which really fast to check but it's right now work only on smallest model and it's not scalable in current state of this direction.
So what we actually use in practice now and how it actually works now in our system uh system verifies not all requests but only some small subset of requests. So each host sample request from another host if it should verify it or not. It for this request repeat the inference to generate exactly same tok same sequence of tokens but with known in advanced sequence of tokens. So it actually make prefill pass and uh get log prop uh log probabilities probabilities to generate tokens for for the same sequence and then it's actually artifact uh which compared to the original artifact and host made decision if it agree if it thinks that uh this inference was valid or invalid and then vote and then exact voting happens only when some host wanted to invalidate another host but it doesn't voting doesn't happens for like successful cases for uh good cases when we think that uh when host thinks that inference is honest uh yeah so when someone when network detects that some inference is invalid hosts which executed this punish pretty hard and it's pretty costly for such host.
Yeah. Uh >> well, what is considered like invalid? >> Uh yeah. So I think for example uh network host Kimmy uh model in F8 uh precision then some host decided to optimize instead FP8 it decided to host in FP4 and it's already not that good quality model and it must be marked as invalid because quality is worse than supposed to be >> but I think uh I'll add a bit here that uh it's actually easy to formulate uh this question in terms of model quality but what the algorithm is checking is not if like the quality of the response is good enough but it's just like whether weights of the model are the same whether uh the output was produced likely produced by the same model.
So it's measuring just the distances between the log props and trying to catch statistically if the host substituted the model or not. If they decide to host more expensive higher precision model, it would still be marked as invalid. Uh that's like not the question that's being checked here. >> It's purely comparing of distributions over tokens with with the same prefix. So yeah. Yeah, you had a question. um for hosts who like call out another host for being invalid.
Um if if they're um if their request for people to like check that the host is invalid is correct like if they're if they successfully call someone out uh do they get a reward for that? >> Uh in current system no uh it's kind of what host must do. So >> yeah, so they are they they participate in epoch and they they have a lot of work which they must do and they are punish and they being punished if they're not doing this job. >> How so?
So so are they assigned to someone to like verify someone to check? >> Yeah. Protocol assigned to assigned such work to verify something. Yeah. >> Yeah. Yeah, it's also randomly sampled and it's the rate of validation can also depend on like the participants behavior like if you are caught cheating more often then you'll be validated more often. Uh what else? There was something else. Uh oh yeah but also about rewards there is um there are rewards generated by the protocol that are distributed based on the weight.
So if you catch a cheater and that host is excluded that like they're not getting the word uh corresponding to their weight, it gets distributed among everybody else. >> I see. It's not really true, but yeah. >> No. Okay. Uh anyway, this is something that uh like part of the protocol that can be changed if the hosts decide to change it and vote to implement their change. >> Uh >> oh. Right. Yeah. >> Yeah. Okay. Uh second part is reliability and so actually we want to again to check that host are not doing some another job on the same GPUs only which protocol assumes.
So the idea is that load is distributed proportionally to weight and host when network is loaded host hosts are fully loaded. they cannot process more uh requests and if so how did it checked it checked how much tokens host produce at time if it responds to inference in time and similar approach and all uh requests which are not processed or correctly processed are also voted by majority it's also decision of many host to mark something as not processed And in the same way as invalid if something is not processed cost is being punished for this and just not getting reward for this ep.
And the latest one is about again scalability. Uh when we started this protocol we didn't really think that it's so important but then uh it was quite painful. Uh yeah so the problem is even if you have like couple thousand GPUs and they are used for inference you need to record millions of transaction to network just for voting just for like billing transactions. So request started, request finished and all the stuff and you everything like that must be not decided by user but decided by the network unless network is decentralized decided by everyone and as turned out no uh decentralized protocol now can process so many transactions.
It's just not work this way. Uh that's why we had to came up with this idea of defarts as we calling this. It's like when user wanted to interact with protocol, wanted to use inference, they must create escrow and open uh small independent uh decentralized network, small independent blockchain which actually lives only for like next couple hours and that's it. And when we have like thousand users, it's uh thousand small temporary blockchains.
Uh yeah, and this can be scaled linearly and it works pretty good. And most important, it's blockchain which only user maintain data for. And yeah, uh just fun fact. Uh yeah, let's go. >> Okay, I'll take over from here. Uh yeah so a bit more uh theoretical kind of um longerlasting research goals is uh what's up with training on this kind of system. Yeah because uh inference you can run on GPUs all over the world you for a lot of things there you don't really need the network.
Yeah, inference would be working, verification, some other things not, but like inference would be running. Now, when we're talking about training, a ton of different problems come up because uh to run a certain model, you would need a cluster of a couple GPUs, but then uh to train the same model, you would normally need to have thousands of GPUs located in the same data center. So, um what kind of problems come up? Yeah, one big part is about training large models and how can we do this not having uh a huge data center, not having this uh ultraast interconnect between GPUs that are sitting in that data center uh and yet achieve uh comparable performance uh and not to lose on the resulting model quality.
So how can we make training over the internet possible? Yeah. Uh big interesting open research question. The other one, similarly to uh inference validation, you also need to validate what's happening with training because you cannot just assume that participants are doing what they claim to be doing. Uh they can be just producing some random steps and submitting them. And here maybe the stakes are even higher because if someone is cheating on inference, you would uh lose on a single inference request.
But if someone is poisoning a huge training run, you might lose the whole run uh and the work of all the other honest participants here. So a bit deeper dive but mostly like overview of research that's being done in this area and kind of how this can move forward. Yeah. So the go distributed training uh yeah the question about how to make uh training of large models feasible in such environment. Now uh there are a lot of gradual steps that have been done in this direction. uh probably the most important foundational one uh is a paper produced by um deep mind uh Google's deep mind group uh about DLCO the distributed low communication training uh which basically uh describes the approach of um running the training in such a way that instead of uh doing all reduce on every training step and uh when GPUs or when all different nodes exchange um gradients on every step of the training.
Uh this this is actually the bottleneck that is requiring um the fast uh interconnect in the data centers. This is the moment when you exchange all that data. This is when you need to transfer uh tons of information. So this is what makes um kind of traditional LLM training transformer training impossible in the distributed environment because if you need to pass that much data over the regular internet it will takes you thousands of years instead of like something that would just take a normal amount of time in a regular data center.
So uh in the decoy paper they explored uh an approach of actually reducing those communication steps and splitting uh the process into two uh into the training loop when you have uh inner optimizer used locally for every participant training the model and much more rare uh steps without optimizer. uh that um that is when you actually aggregate uh the steps that everybody computed separately. So you're parallelizing the training.
You're still having this uh steps covering everybody. uh and the results are showing that the quality of the models produced in such a way are actually hitting the benchmarks of the models of the similar size uh trained in a centralized colloccated way. Uh some mentions but uh not the full list of uh training runs done in this direction. Prime Intellect had I think the kind of first uh open verification of that uh protocol working.
Um one of Banzer uh Bananzer subnets trained a 72 billion parameter model in this way. And then also like some pretty cool research from Plurales group who also showed recently that uh not only you can do the um training over the internet in terms of speed of connection but you can also uh solve this problem for memory and uh they actually managed to feed the training run on distributed and heterogeneous consumer grade including hardware.
So that's one thing that's developing pretty quickly which I think is also very interesting not only for decentralized networks but pretty much like for everybody who is does not have resources to build a gigantic data center that's bigger than every predecessor before them. Uh what's next? Uh the other thing is yeah oh yeah I had a question on dlo does this assume that all of your diff separate kernel workers are about like similarly fast um or that they can achieve a similar number of iterations in the same amount of time because I'd assume one challenge of this is if you have one worker that's like half as slow as the other workers and your outer optimizer assumes that all of your all of your input streams are like are synchronized at the same time are to the same number of steps that like would that be an issue or >> uh yeah I believe that the dela runs were done on the homogeneous same hardware uh heterogeneous hardware is a completely like different piece of that pro problem that doesn't come for free that's why yeah I mentioned pure is wrong because it's like I think the fun yeah >> fact to mention here that they kind of investigated two cases first if all your workers process the same data the same data sets but many times or if different workers go over significantly different part of data.
So it's actually similar what you're talking about because you can have some data which can contribute a lot to the model and you can have some not that good quality data which not contribute a lot and I think it's quite well aligned uh with approach when some worker is slower and so it's actually works not that fast so converge not this fast but it still converge what is important I think mostly because they still doing this all reduce pretty often.
So it's like one in couple hours they still doing this all reuse and still synchronize results from all workers. Yeah, >> that makes sense. Sorry, I I I think I misinterpreted heterogeneous hardware earlier to mean like um over different outer optimizer steps you could have different sets of you could have different sets of inner parallel workers. Um does um are there any methods where like over different optimizer steps you could have like varying sets um what is it of various sets of of workers in in your inner layers like for example say you have some person who might want to use their allocate their GPU to this training for like some not for like only a day for example but you want to train for like 3 days for example so they can only contribute for some number of the outer loop optimizations >> uh it's actually much simpler and even so uh open da is what prime intellect uh worked on a couple years ago.
They just reproduce the paper and they allowed uh to join and uh cancel your participation in this training and it's actually pretty simple. So you just participated in some outer loop, you disconnect it, you not contribute to another after each or reduce after each uh outer step, it's not important who participate in the next one. >> Oh, I see. Do you have to um does like your optimizer have to like uh have to like account for this?
Like if I don't know what the outer optimizer exactly, but if it's something similar to atom or something like that. >> Yes, it's >> would that create some instability if like you're losing u you're losing like parallel worker or like would that make your would you have to do some adjustment to your optimization? >> Uh I cannot like answer you exactly. I know that uh it's part what they covered and it worked but it would be I I cannot answer exactly how they uh what they do when it's uh cancel the participation but as far as I know there's almost no custom logic there.
It's just like pretty pure from torch import Adamw and that's it. Yeah. That's um yeah also kind of comes as a somewhat given requirement because you cannot guarantee that a participant would stay in the training but uh in in a sense yeah if you're training in a data center you have no control over a GPU just uh dropping out because uh it burned a >> uh it may be also fun to mention here how covenant AI works so covenant 72 billion models so they kind of didn't use so they added one more thing to account for untrusted environment here uh at contribution of each worker in outer loop they first estimate the quality of this model for some uh test set and so and then when they aggregate uh contribution of different workers they kind of waited it >> uh from their opinion how they used it not really promising but they they made this attempt and model itself was pretty successful.
Uh but the idea to weight results based on some benchmarks based on some tests it's also sounds pretty good from my perspective. >> Oh I see. Also, is the intuition for like the outer optimization similar to is it similar to like um to like model souping um where you're just like averaging the different weights together but waiting by something or is it more of like or like is is the if the out if the outer optimizer is just like atom or something you just treat each of the different parallel runs as giving you some like as giving you some like delta w for each of your weights >> it giving you delta w for each of weights it's kind of uh similar to aggregating models But it's important that diff between these models are not that significant.
It's still incremental step from different uh contributors. It's still like training only for this last hour. >> I see. But but but you're not looking at like the at like the time series of delta w. You're not looking at like the gradients from every it from every single step in the lab. >> You're not looking. Yeah. >> I see. >> Yeah. Uh but yeah the big difference here also is that in case of log everybody host the whole model in the memory that's not necessarily like a requirement that all how do like it gives you a limit on the lower limit on who can participate.
Uh but yeah uh the fall tolerance problem and GPUs dropping out doesn't go away regardless of whether you you store the whole model or not. Uh that's a lot of yeah a lot of the interesting challenges but yeah uh the trustless part um the trustless thing didn't go away and we still are in the system when you cannot uh necessarily assume that everybody is doing the work that they're claiming to do. So how do you make sure that um your whole training run doesn't go to waste? uh it's not the question about just catching and eliminating participants that are trying to poison the model or that are you know by mistake just submitting some um weird results.
Uh but here the goal is yeah not not to just find those participates or not to avoid some weird incorrect submissions completely because that would also be impossible. Uh but the goal is yeah to make sure that eventually your model would converge uh to what it would converge in the honest case regardless. Uh now the requirements you'd want to have in this case yeah is that even with some percentage of uh potentiously malicious or just you know randomly incorrect uh incorrect updates the honest ones would still continue to improve the model and get it towards that conversions. the at the same time the faulty updates uh would not dominate on that process of averaging uh aggregating and whenever uh yeah whenever you're aggregating the results from participants uh and then you obviously yeah want to run verification um try to validate the steps of what every partic case if computed uh but you don't want to add with this verification as much overhead that it would make the whole run impossible time wise.
So you need to verify what people are doing but not too often so that you still don't damage performance significantly uh yet well enough and often enough so that the model would converge to what it's supposed to. Oh >> yeah, it's also fun to mention here that this problem of like so model keep improving. So it's uh bizantine tolerant training. It's kind of right now approached from research community more from theoretical perspective.
There are a lot of like real math about this and real and then based on this theorems based on this proofs they do experiment only after. So this uh paper of Edward which is mentioned here it's approach the problem from this side that's pretty funny because it's not like standard approach for what we see in uh AI world and yeah >> yeah it's yeah it's it's indeed kind of like a theoretical setup but also if you account for the real world then uh there can be you know uh uh you don't uh just think about like oh can we catch and eliminate You also understand that again rewards and any other incent like any incentives that uh are happening in the system would actually affect uh the behavior of participants and uh affect whether it makes sense for them to cheat or not.
Uh yeah so uh a bit uh like different formulations come to the problem uh given you know the real world constraints. Um what else? Okay last uh mention beyond the training. Yeah, that uh we just spoke about training large models and trying to compete with centralized data centers in the size of what you can train. Uh but we can ask a different question. Yeah. Can we actually train small models and make them communicate between each other so that they would large uh they would match the quality of much larger models?
Uh the first mention here that's worth uh taking a look at is Deepacco same group as uh introduced deco uh proposed uh the way of scaling training runs by training uh independent modules of the same model. So that that that that kind of like different approach from diloc that we just discussed when you would aggregate all the time. Yeah. Here you you separate uh completely. Now uh another interesting research approach that shows very promising results is about latent uh collaboration and that is when you have smaller models that would uh exchange instead of um a collaboration when the models would talk to each other uh in like the same reasoning steps and kind of plain English text as you would have like agents run right now.
What if the models instead um exchange their hidden states? Uh and it turns out that uh on many benchmarks uh it shows like that pretty much like half of the difference can be cut uh between the model sizes when combining smaller size models. No, they actually won as a >> Oh, yeah. A >> I forgot name. >> So the startup magic >> AJ benchmark on kegle. They won it like couple months ago. Uh with like using for decode step model which like 100 times smaller than second place something like that but used bigger model for prefill steps.
So that sounds pretty good cool because it's really cheap to compute this way >> and all these things are important for decentralized world because once there is good results in both training of such smaller models which interact and inference of such models you can really utilize not only like data center grade hardware uh to work in decentralized network but also you really can utilize your phone. So if uh one billion model will be really valuable when interact with thousands of another 1 billion model uh such compute of cool things can happen in our in our pockets and yeah so I think it's one of big shiny goal how it can be truly decentralized.
Yeah. >> Yeah. >> Yeah. five. >> That's it for now. >> Yeah, probably that's it for today. Uh thank you all. Uh I just wanted to add that most topics we touched today are not really solved ones is just some uh problems which such system has and such systems have because there are many different system like that and they all still open in in none of them. there are real like final results with which already good enough for long-term use in production.
It's all pretty early in development and in all of them there are many both engineering and research questions and yeah we will be happy if some of you will decide to to help uh >> I take a look into this again uh in in everything that we mentioned I think there is like some working step uh it's not all like completely hypothetical inferences running uh training runs that I mentioned happened uh yet I think on each of those directions there is a lot of uh new stuff that can be done uh if you want to reach out there are last minute uh uh Twitter links but also don't uh hesitate to just check out the GitHub of the projects skunkai and uh like browse it's it's fully open source so you can uh see and check everything for yourself and uh yeah happy to keep chatting happy to answer any questions Yeah, I guess I'm curious.
Does that have any like data retention guarantees around like like if you're submitting like infos or like training jobs like wouldn't it be a worry that like some random like person offering up cookie could just say do you >> uh not random person uh but definitely right now in current uh approach how it works uh the server which execute your inference can have your data. So in their RAM it will be available for host uh and they can if they want to get access to it and also another host if they must validate inference request they also can access to your data to make to make the validation because they need to repeat it.
Uh but it's only host no one another in the system. So we for example we are not really hosts we are like uh building this we don't have access to your data and so it works for many users but we clearly understand that such requirements would not work for the big corporations and it's kind of limit when who can use such system and but it's not unsolvable problem and there are pretty big and promising thing in this direction is trusted execution environment systems.
So when actually hardware itself guarantee that everything which is in RAM of the server is encrypted and even if you have physical access to the server uh you cannot get access to data and that's right now doesn't work for such system because not all hardware uh has this technology but many new have and we hope that it will be supported on chain pretty soon. Uh but yeah, here this topic kind of even bigger because it requires also a lot of research in how hardware works because right now it's quite often new papers about oh we got couple million dollars we got our own laboratory and broke this system.
It happens. Uh yeah >> I think two maybe things to add here. One is on the guarantees of data retention that you have um with other say open source model providers or even closed source. Yeah. Like in most of the cases the host that is providing you inference can access uh what's run there. Uh there are separate again options supporting TE but they often also marked. So it's not something that's like happening elsewhere and only decentralized system would not uh completely vault your data.
So that's one thing to consider and like some part of this production would be just done uh as a form of like some agreement that you are getting when you are getting the services but on the other hand yeah >> I think uh it's good point to mention about agreement. Yeah, the one approach how we it's not implemented but how we consider uh it can be solved also. So there are quite uh a bit huge hosts hosts which uh publicly known which are companies with reputation and even in decentralized system they totally can uh sign the same agreement that they will delete uh your data in some period for example in 24 hours and if they sign such thing you can choose to schedule your requests only for to such host And then you actually will have exactly the same guarantees about your data as they are now for example in antrophic on openai.
Yeah. >> Yeah. So theoretically that's like not that much uh different from any other provider. The other thing about trusted execution environment, Gleb correctly mentioned that it does work on all hardware and uh like my my point of view on this is that um a trusted execution environment can actually like solve a couple problems for you. Yeah, it would solve validation for sure. um it will you know you you won't need to revalidate what they did if they did it with a closed signed container and you can verify that exactly that work was completed and it will solve the data privacy concerns.
Uh in the same time uh it would be prohibitive for hardware uh that's uh doesn't support this technology to join. So I feel like it's very important to uh still continue working on uh validation verification approaches and yet to make sure that whoever can support this technology can turn it off and can like uh offer it as an option. Oh, could you also speak more about how um for training how you would deal with um what with like adversarial um of parallel inner workers um or how you would be like or how you would be robust to that?
Uh what do you mean >> for what is it for for further distributed training like for I think it was um um um dilaco um you mentioned that um you're you've designed the training pipeline in such a way that even if like percentage of the workers are like malicious and are beating you like bad updates you will still not be like like you the training will still like overall work or you'll >> I think here it's important to mention that we talked about not only what we tried and we developed but only about another approach.
So the Byzantine tolerant robust training is not what we worked on. It's just >> something that definitely needs to be done yet on the network there is still no robust mechanism for uh adver adversarial participants in the training. Uh but >> during test net we had like a proof of learning approach is actually when so we have workers u many of them just doing gradient steps in parallel on different amount of data but we seeded how in which order they process data and we enforce them to dump gradient step for some certain random steps and at some point another host can uh request and just check the the single gradient step.
The problem of this approach that it's extremely expensive in terms of data saved on disk because like such dumps are huge. >> Uh but it allows like pretty fast to detect who is not doing work. >> I think it's theoretically awesome but not practical at all. Uh yeah. >> Yeah. Uh I saw on your website there was like cost per imperence of like some of the popular open source models and stuff. Uh how are you able to like determine that like mix that in so is that kind of determined by like how many GPU providers are even though >> uh >> is that cost difference?
Yeah, >> it's how many user pays but uh actually again the system design about rewards distributed per uh per epoch and it has two types of incentive. First one is just flat reward which is uh minted in this epoch and distributed in the end and second part is really how many user pays to the uh to the for for the inference and the tokconomics of the system design that first part is much bigger in the first years of network life than the second part.
What I mean that right now it's crazy cheap for uh users and they are not really paying to paying enough to fully pay for the hardware of hosts and right now hosts are getting paid from the from the protocol itself for the fact that they are contributing hardware I think also important to mention so more like approach which we had when we launched but right now it's kind of post-temporary it's ideal of dynamic pricing.
So to balance how many hosts and how many loads from user on this on them uh there was idea uh again which right now it's not on chain uh that pricing will be dynamic if uh the utilization of network insert certain window for example from 40% to 70% the price is flat and cheaper than marketplace price but if price is becoming if low utilization is becoming higher. It's already network would not be able to uh process all this request with good quality.
The price will be will become a little bit higher more load more higher and then it will create incentive for user to just use network less and uh and and and utilization will stay in like the target window. But in opposite if the network is not utilized price will just drop lower and it will attract more user to the network. >> Yeah because I guess like you know it's bad for everybody if there's like not enough utilization and yeah >> price is still high. >> Yeah question I had is like how do you make sure that people are like detecting that they're being validated I guess like I don't know >> uh inference >> inference.
Yeah. >> Yeah. uh they kind of first have all commitment to the artifacts recorded and only then it's sampled if this uh inference will be uh validated or not. So yeah >> there's kind of like probability with which you would get validated but >> I think important question that it's it happens after all something happens after the something >> yeah after it's generated. Um would there also be a way for inference or like for for distributed inference to be compliant with um with any like regulations?
Like for example, if the government uh makes a law saying that if you're running inference, you have to like you have to have mechanisms to block requests about like bioteterrorism or something like that. um like would there be a way to enforce that different inference providers are abiding by any um by any like set of rules or regulations? >> Uh I think it's important that there is no like it's not single inference provider.
It's uh a lot of independent hosts in a lot of independent in a lot of countries and they all have their own regul regulations. So uh it's kind of what host supposed to solve and also if someone resell inference from Gonka for example in United States they also must for sure add such layer of detecting such regulations but the idea itself that it's network to decide which requirements exactly will be on this network and if some uh if network wants to adopt such regulations they want to uh hosts must vote and uh update software and update models to satisfy such requirements. >> I see. >> Yeah.
So it's possible but it's only up to network to decide how exactly do it and it's also must be like governance based like network governance I mean voting by host based software update. >> That makes sense. Wouldn't do you think some of the verification technology being here being being used here could also be repurposed for like say if the government wants a way of like um having like a verification system that different providers are like abiding by like a specific set of rules.
In this case, you're trying to verify that they are abiding by the rule of like you are running like the actual model that you say you're running. Um would there be a way for this to be like repurposed for like for regulation enforcement or for validation like that or would this >> like >> I actually don't think it can be because the system of validation don't know know exactly nothing about the content of request so it doesn't understand the topic it doesn't understand what it's about from my opinion it works only in terms of uh again distributions which model which model has on certain prefix and And it it can be used only to check if it's really certain model. >> I think I'm not sure if I understood your question.
If yeah, if you're asking about uh like content of request or something like this, then it's a completely different thing. If the question was about the validating if the model was like model weights were not changed, were not altered or something. Yeah. Then I mean yeah, sure it's open source. It's like >> or I guess like my initial question was specifically about like specific user requests, but like I guess like could you also imagine a case where um where the government makes some regulation that you have to like um you have to include like a specific steering vector or something like this into your model to like keep it like if if there if you wanted to enforce a rule at like the activation level at the model architecture level that prevents and answering like bad requests, could this be used for that?
Maybe >> I think but again for the host to decide there is it it doesn't work the way that >> technology it's not I mean okay >> yeah so for example if governance enforce it it can then enforce company to provide some artifacts of how this model with version a uh behave on certain requests and then this technology can be used to like check on similar request or something like that but I think it would require a lot of experiments to to understand properly if it could solve problem or not. >> Any question? >> Yeah, I think if there aren't any more questions then we can just wrap it up if you guys want. >> Last question.
How'd you guys come up with the name? >> Oh, it's actually uh we are Russian speaking. uh there are a lot of Russian speaking people in this company and this name it means like >> a race >> race yeah who run faster and the idea of uh this proof of compute thing is to like race in optimizing your hardware so you're optimizing hardware for the model for the serving and yeah that's actually how we came up with this >> thank you thank
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.
Filler phrases
803 in total: uh 420 · like 223 · um 60 · actually 53 · kind of 32 · you know 7 · I mean 4 · basically 3 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.