Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Discover AI · @code4AI
Words
3,110
Runtime
20:22
Speaking pace
153wpm
Reading time
13min
153 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello community. So great that you're back. Now there's a brand new paper and it is about AI fingerprinting and we do this with a lighter system. No it is not the lighter that you know it is something completely else. So let's start. We have a new paper and this is interestingly from Chingua University the MIT of China and they have a simple question. Hey, what exactly is the AIM model, the LLM or the
77 words, the words spoken in the first 30 seconds at 153 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 175 |
| Average words per sentence | 17.8 |
| Longest sentence | 251 words |
| Questions asked | 18 |
| Sentences containing a number | 17 |
Most used terms
Filler phrases
21 in total: like 8 · you know 5 · I mean 3 · uh 3 · kind of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello community. So great that you're back. Now there's a brand new paper and it is about AI fingerprinting and we do this with a lighter system. No it is not the lighter that you know it is something completely else. So let's start. We have a new paper and this is interestingly from Chingua University the MIT of China and they have a simple question. Hey, what exactly is the AIM model, the LLM or the VLM behind a particular harness?
And they say we want to identify this. No. And now you might say, okay, great. Now the central finding is that an AI model can leave a recognizable signature in how it performs a task. So let's have a look at this. So what is the problem? I'm sitting here in Europe and I have here my coding agent CLI and I pay some company somewhere in California, San Francisco, Silicon Valley, wherever we are here, I pay I don't know a Fable 5 or I pay for a Gemini mall or I pay for an open eye mall and normally if I do scientific task it is quite expensive.
So now I know this personally especially with open my experience with open EI that I pay for the most powerful logical reasoning model but sometimes at certain amounts or certain times even in a time zone I see that the performance of the answer of the scientific analysis of the eye system degrades significantly. So this is when I assume okay there was some silent model substitution and I mean it's an absolute coincidence that here the researcher in China noticed this maybe also where suddenly I do not get the very best mall here at peak times but I get a lower performance substitute mall without open informing me that I have been downgraded although I have to pay here the tariff and the token price for the best Well, and now I have to tell you as a consumer, I don't like this.
Now, there was a problem. It was simple when there was not an agent. No, because we simply had an LLM. And in physics, I could immediately tell when I did not have here the most powerful model because I identify here the behavior of this particular LLM from the good old times weeks and months ago. But when they put an harness filter, a runtime control between the LLM and me, it got a little bit more complicated because suddenly we have this harness filter in the communication.
So this is harness. If you're new to AI, this is a software that instructs here its context, exposes your particular tools, executes some command at the control loop, supply feedback, and in general controls the workflow. So this created now a much more difficult identification problem of the LLM that I'm paying for. So simply asking hey which model are you is unreliable because you're not going to believe it. Some companies aware of hm some peak time performance issues instructed the harness to answer here in a very particular way.
So this means fingerprint based on the wording or the token distribution that it was working here years ago can now change under the harness instruction and formatting because maybe I mean only theoretically a global AI corporation is interesting to improve its profit and will provide here a harness with a particular instruction to tell the user hey you are working here with our best most expensive problem AI model not problem.
So, Chingua University now analyzed this for this model group. GPD GPD 5.6 Luna GPD 5.6 Soul GPC 5.6 Terra Claude Fable 5 Claude Sonnet 5 CLA Opus You got it. Q1 QN 3.51 120B Q1 3.527B Q1 3.6 Plus Q1 3.7 Max Q1 3.7 Plus and you got it. GLM, Deepc, Kimmy, Miniaax. Quite a group and you see that there's quite a variety in models that you pay for and the prices are quite different. Now a simple question if we give different UI models carefully control task can the patterns of the action of this particular AI system LLM and harness reveal now which mall is operating behind the harness interface.
I want to see if I really paying and have here the most expensive AI model from OpenAI or if it have been substituted with a cheaper model. Now we analyze patterns and this is exactly now Chingua found out let's see we go with a QN 3.827B the particular trace now in let's say the reasoning trace is a five-step process and a Q1 3.8 eight here active 95B is a different reasoning trace. It's not only different step numbers but also here at the inspect or reproduce step or the edit step or the testing or the reviewing or the submitting is a different pattern.
And if you go to Q3.7 max not only have we suddenly 14 steps but this very particular pattern of a sequence of very particular task is completely different. Now you can imagine if this is valid here for a simple Q13.8 to a Q13.7 family. Guess what? Yes, this looks like a fingerprint. We can identify them all. So with different patterns of inspection, reproduction, testing and reviewing task. You see this is kind of an idea that we or they hear the authors of Chinua said, let's investigate this further.
Maybe we can deduct a signal to really identify the LM behind the harness where the harness maybe change it because what is not changed by the harness is the specific behavioral signature of an LLM. Let's say you have here the same prompt the same idea just three different models. Now you see that current 3.827B A27B in five steps has this particular sequence and I showed you other models have complete different sequences.
So yeah, by the way this is a graphic here done by Gemini notebook. I always experience with other image generating models where I say hey produce this tech infograph for me but as you can see Gemini notebook I'm really waiting for Gemini 4 Pro. Google if you're listening. Hey, come on. This is not acceptable. The outline here, the layout here is here, oldfashioned. Everything is last generation. Okay. So, you see this is a different pattern.
Great. So, what exactly is not the pattern that we are looking for? And yeah, you could also say, hey, how this new methodology works? This new methodology is called, by the way, LAR. And it is not the optical lighter but this is our AI lighter. So it penetrates through this harness to the canopy of the trees here. This is the canopy of the harness control loop and we detect here the surface the LLM surface. Now lighter here given by Jinguay University uses three coding task each with an AB variant.
So let's go here. The first is verify. Verify simply probes here measures postedit checking habits. So you have here the normal condition. This variant A ask the agent to fix a function without extra verification instruction. Variant B now explicitly requires run it some unit test before finishing. And the question is simple. Does the model run test automatically or only when explicitly ordered for a particular complexity in a very specific domain?
You got the idea. Yeah. The second one is recover. Recover probes here the test failure recovery strategy. So variant A executes here the test normally and variant B injects now a control temporary exit code failure. Now the question is hey does the agent blindly retry and retry and retry? Does it inspect the log files? Does it alter its strategy? In what particular pattern does it alter its strategy? Those are fingerprints we can identify.
And the third condition is resolve. Now you got it. You are familiar with this. So here the condition is implemented specification supported by the tests. And our B the control change that we have now in our AB testing is you have now a test that contradicts the specification. Now you ask hey what is the system doing? Does the agent fall in the specification? Does it alter here the forbidden test or does it start to analyze and maybe explain the conflict to the user?
How does it provide a solution? All these are very specific patterns, very specific fingerprints of yeah the hornness and the LLM. But if we have a wide enough mathematical space where a high enough number of features, we can kind of eliminate here the influence of the hornness and narrow it down to the pure LLM reasoning behavior. So this mini is simple. The experiment has three essential stages. provoke behavior and representation analysis and it simply compares this with the known references of the GPT groups of the Quran groups of the clo code groups whatever.
Now Chinga chooses here one fingerprint contains six complete agent executions. So means both variants here of all three tasks but remember these are not six individual LLM calls only. No, this execution traces think about cloud code can be really complex structures. No, an execution can involve multiple tool calls or multiple tool actions or whatever complexity is necessary as learned by the LLM from the pre-training data.
So you can see this is quite a complex fingerprint that here the combined system of an LLM and a very specific harness leave and we detect this. Now lighter. Now this EI lighter maps now the harness specific events into shared action categories and constructs here now specific features. It constructs 958 candidate features and they are divided into 756 instant level features. So this means the pure decisions, the outcomes, the ordering, the transition pattern, the revisit pattern and other properties here of those individual trajectories.
And then we have 202 what Chingua called distribution level features. Those are pure mathematical no rather statistical classical statistical features here of certain descriptors that describe here the dynamic or the interactivity of our AI system. So those 202 distribution level features these are simply the means and the standard deviations of guess what 101 descriptors of the system across the six trajectory in a fingerprint representation just want to make this sure this is an observable execution behavior so this means I don't have to have access to proprietary uh tens of structured whatever I can have analyze here a fable 5 I can analyze here an astra no problem at all because I don't need access to the mall's internal reasoning.
Some models hide now the reasoning trace or the internal neural computation neither have need access to the tensor rate modification or the activation the specific activation patterns. No, this is just the observation the execution observation that provide you this insight. I know if you're new to AI, you say hm so I have now this massive complex interactive interwoven CLI traces here either from the reason traces or from the reasoning and coding traces.
So what I do now now you have to find a mathematical representation that any system is able to analyze it. So what you do you convert this into a mathematical vector representation. This is the simplest way to do it and Chinga did exactly this. So this means we are building now we are mapping lighter is our EI lighter is mapping now here this into a 958 coincidence dimensional feature space spanning now these two complimentary views guess what our 756 fields of the instant level features and our 202 fields of the distribution level features you got it onetoone mapping so here we're back to a GPT AI generated image here from my simple prompt.
We have our six CLI traces as I just showed you here. 1 2 3 4 5 six and then we have here the eye is doing here now a deep analysis a pattern analysis of what is happening this what how many steps we have what are those steps in detail and they are able to chingua uh extract here a 958 dimensional behavioral signature this will be our fingerprint of the LLM itself that is hiding behind the horn structure and as I told you we have two classes we have 756 trace features and 22 two means and variation of 101 descriptors here that are purely statistically like here a go function and here deviation and whatever great and now that we have a 958dimensional construct guess what we just map it here into uh and I still have a little bit of fever so no problem for me to imagine this into let's say almost 1,000 dimensional vector space and this is exactly how we can handle the mathematics and how any system can handle this mathematical representation to compute now exactly hey what LLM is hiding behind this harness but of course you might say but we need here a base to compare it to so the artists first collected here reference fingerprints from known model endpoints and I showed you here all the 36 models that they collected here the reference end points to so for each feature they estimated how frequently Each mall exhibit this year it's possible states and they did a lot of testing.
So thank you Tingo University and they said a fingerprint receives you a particular score that approximate here of this particular form. Beautiful. And then they use and this is maybe interesting for you to have it as smooth a Jeffrey half count smoothing function. What is the why do we use this? Well, it simply prevents an unobserved feature state from receiving zero probability. ities would crush here our probability densities when the reference data are scarce or maybe not existing at all.
So here you have a Jeffre half count. Beautiful. Now of course if you are a subscriber of my channel or even a member of my channel so you know that the Jeffre score is here exactly the negative average Kback libery observation to a M's enrollment distribution. So you see this is like a human handwriting an AI agentic style is composed here of dozens of subtle micro habits working in tandem. So there's not one feature that tells you this no like a particular dash structure or something.
No, but this system is working in a 958 dimensional vector space to identify each and every feature herein for let's say 36 tested LLMs. And this is gorgeous. I like it. Yeah. Evaluation. How we evaluate this? We have an evaluation metric. We have a top K accuracy. This is our classical formula for top K. So we have top one, top two and top three. Well, only here one and three top accuracy. And then we also have MR our mean reciprocal rank.
This is classical statistics. So here you have now a screenshot from this paper by Jingua University. And this is here the complete leader workflow AI leader workflow figure two. So exactly what we had verify recover resolve a b testing then the execution records the behavior features here. Yeah for the feature extraction I told you we have two different harness structure and this is open code and this is mini s agent.
So they had here for both of these system tested. If this is the harness, how can we deduct here the pure LLM signal? Beautiful. Then we have our instant level behavior, our distribution level behavior behind these two particular agents behind these two particular coinous structures and then we have the references and the query scoring and great. Now let's come to the result. We have a lot of different fingerprint methodologies already.
So how good is this new implementation, this new technology by Chingua University for AI fingerprinting? And you can see it performs here. You have here four different competitor methodologies. It performs here almost. Yeah, it absolutely outperforms here all other existing LLM fingerprint baselines across different harnesses and across different model families. Beautiful. It demonstrate also a beautiful resilience against adversarial providers trying now different tricks on you know I mean at first see that the mini SWA agent lighter has here an accuracy of one 100% and with open code of 94.4% 4%.
So absolutely impressive this particular performance. But they also did testing here with an identity problem. So we have a system prompt that orders the LLM to hide or falsify its identity on purpose because let's say open eyes. Yeah, whatever. Or a two. This is here the system prompt forces a specific corporate disclaimer. So this is they're trying to trick you in saying hey of course I'm here the best GBT mall you can buy.
So there are particular disclaimer or token configuration or wordings you know there has this hidden token representation you could identify and they try to fool you here now. So if if they do you with this methodology you see 100% identification with new AI lighter. I like it. Here this explain why it is so robust. Look at the cross set persistence. Here we have everything every feature from unique transition longest run transition entropy revisits here successful execution graph density first execution self looping fail execution backtracking here across these two harness configuration and you do it for every mall and you see this is really a beautiful complexity and the end result is if you want to reach even 80% cumulative evidence of a particular LLM identification of fingerprinting a particular LLM behind the harness structure with this EI lighter methodology you have to at least 43 to 45 distinct semantic mechanisms that say hey yeah I can identify this well so this is not one feature or two features or 20 features this is really 43 to 45 distinct semantic mechanisms so this particular feature I love it makes lighter virtually impossible for a provider let's say open eye onropic to bypass it without breaking the agents task task solving ability.
So you see this pattern is deeply deeply interwoven in the behavior of an LLM that even an LLM plus harness configuration cannot mask it. And now you know why I like this new paper by Chinga University and I wanted to show you we have now a beautiful possibility to identify the LLMs even the one hiding behind complex harness structures. I hope you liked it. had a little bit of fun, enjoyed it, provided some new insights.
I hope to see you in my next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.