Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

JakSec · @JakSec
Words
2,737
Runtime
13:53
Speaking pace
197wpm
Reading time
11min
197 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
So, last week I covered more or less the foundations of regex for security, some of the more common mistakes that can happen, you know, how regex works, and one example from Ruby that could actually have that to RCE with regex. But, today I want to look at some actual write-ups from some bug bounty hunters who have, you know, looked at these regex vulnerabilities and identified where their weaknesses are. So, in the way that the developer writes them, some of the assumptions that they made. This is obviously a natural progression, but let's have a look
99 words, the words spoken in the first 30 seconds at 197 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 135 |
| Average words per sentence | 20.3 |
| Longest sentence | 72 words |
| Questions asked | 8 |
| Sentences containing a number | 3 |
Most used terms
Filler phrases
153 in total: like 49 · you know 39 · basically 22 · kind of 19 · actually 14 · literally 4 · right? 3 · uh 2 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
So, last week I covered more or less the foundations of regex for security, some of the more common mistakes that can happen, you know, how regex works, and one example from Ruby that could actually have that to RCE with regex. But, today I want to look at some actual write-ups from some bug bounty hunters who have, you know, looked at these regex vulnerabilities and identified where their weaknesses are. So, in the way that the developer writes them, some of the assumptions that they made.
This is obviously a natural progression, but let's have a look at our real write-ups, see how these vulnerabilities are going to exist in the real world, which is sometimes different than, you know, theoretical stuff we do today. So, the first vulnerability is going to be a blind SSRF. The second one is going to be an OOB vulnerability, and the last one is going to be a general XSS write-up with weak regex is actually in HTML validator.
So, some pretty cool reports, and we'll jump straight into them now. So, the first report is from this weird Go library, >> [snorts] >> and it's just a basic, simple blind SSRF. I wanted to start with something not that difficult so that we can understand basically what this is going to look like, and it allowed us to basically hit like, you know, internal services, but it was a blind request, and it was a blind get request, so it wasn't really that impactful, but we can still have a look and see where the regex went wrong.
This is one of the functions here that used some of the regexes here to basically make a check, and the regex looked like this. So, you see the regex pattern here. And, you know, if you remember from my last video on covering the basics of regex, you should definitely go watch it, but there's, you know, a lot of common foot trips here with the regexes, and, you know, some of these characters aren't literal characters, but they actually mean something different in regex.
So, if we just follow along with what's going on here, then the regex starts with https {slash}{slash}, you know, there's some domain, so this looks like Amazon EKS domain, so some kind of cluster, Kubernetes cluster, whatever. Then we have like any number of characters, lowercase, hyphens are allowed, underscores are allowed, characters A to Z, uppercase, lowercase, and numbers at least once, but can be any amount of times.
And then amazonaws.com and then this at ID and basically the same thing again, just this time for the path. So, what this regex is doing is pinning the URL parameter, suppose it's some URL parameter or whatever, down to a specific domain with some kind of wildcard in the middle of it, basically, followed by a path, a specific path on that domain, and then, you know, pinned down to only basically ID. So, you you can't go any other paths apart from that.
Now, can you spot the issue? Well, the issue is that it hasn't been anchored. So, anchoring a regex is where we put either one of these characters. So, this arrow character [laughter] or this dollar sign character that will tell regex that this should only be matched at the start of the string. If this is not there, then literally any part of the string will match it correctly. And in this case, the attacker was able to match it with something like this.
Because regex matches anywhere in the string unless you tell it to. It'll match any part of the substring, which is what it's designed to do, and it will return it to exactly as you want it to. And this can lead to a lot of problems. And obviously this led to this validator being bypassed, so there was a SSRF vulnerability here because the regex was bypassed. That was pretty cool. And um yeah, basically in in this service, it led to blind SSRF.
Now, this is only like a get, and it also was blind, so there was no way to read arbitrary like responses or whatever. But either way, it's still a cool vulnerability and shows you just how easy it is for forgetting one character to basically mess up your regex. Next report is in this thing called Authentic, which I think is like an identity provider, as far as I know. But, anyway, they have like an OAuth 2 system, as literally every app nowadays does.
You can register an app and like sign up or whatever. And the researcher going into an OAuth 2 is I assume you have some kind of understanding of it, but basically, if you don't, of course, is one of the apps registers, that you get a link back, and then you approve it, and then a code gets sent back to the app, and the app can exchange the code with the resource provider for your token. If that didn't cover it for you, then here's a diagram for you, basically.
So, you get a link, you go back, and then the app exchanges, you get token back, and now the app can access your data. And the vulnerability was here, basically. So, when you're typing in these redirect URLs, on some websites, you can use a regex with it, right? So, on some websites, you can use a regex with it. And you can type in like a regex pattern, because in some cases, you might need to allow wildcard origins for your redirect URLs.
Basically, the main vulnerabilities around OAuth are attacking this redirect URI, because it kind of leads to these like phishing scenarios, where basically, you entice someone to click on a link where the redirect URI is something that you control. They like authorize your app on like, you know, a legitimate-looking page controlled by the main application, but the code gets sent back to you. So, that's like the main attack surface for OAuth.
And when this redirect URI is parsed incorrectly, that can lead to some pretty severe vulnerabilities, at least in the eyes of the program managers. So, they were using this really annoying Python code. Uh they were using Python code, and they were using full match. So, they were you allowed regex, basically, in this atom. But, what this research I noticed is that most apps forgot to escape the dot on the regex pattern.
So, most people just didn't even know that this was using regex, and they literally just put down like a whole application here and put the dots in. So, if you know about regex, the dots are any character, right? Dots can be any character. So, that can lead to, as they documented, stuff like this where an attacker could register a domain and pull any character except the dot, like an A or whatever be zero in this case, and it would still match the regex pattern.
So, that's where the code would get sent to at the end. And they framed this as the vulnerability. Now, they call this intended functionality apparently because that's just classic code that they always do. But, then he actually found out that like the provider program was like using they had like a cloud platform that was using like this service authentic on the back end. And then >> [laughter] >> And then it it turns out that they actually made this mistake.
And I think they had like a pretty cool video here I'll show to you that they actually like literally made the exact same mistake where he was able to register a domain with a dot inside of it, like an unescaped regex pattern. And then once he went to it, he ended up being able to basically take the session ID from it. And I think as well because yeah, he was able to log in and then yeah, basically the session cookie went into his thing and he had the CSRF token as well.
So, I think that's pretty funny and >> [laughter] >> and it turns out that they they changed their mind after that. So, I think that does show you that you you have to be careful with regex definitely because if you don't tell people that, you know, this feature is using regex, they're going to make mistakes. And I would definitely won't trust users with building good regex because it's hard enough to do as an app maintainer.
Trusting users to do that seems like a pretty bad idea. So, I think the fix was that just escaped these things automatically now for you and it's just probably better that way to be honest. And yeah, so those are the kind of things that can happen with these redirect URIs, especially with OAuth. So, you know, keep an eye out for these unescaped regex patterns. This was a obviously a white box analysis, so he had the source code here.
That's not what is always going to happen, but if you know, register like an app on an OAuth provider, you know, check if the regex maybe there's like a regex pattern doing it maybe the dots aren't escaped. So, it's something worth having a look for as well. And the next report is covering a kind of different class of regex. So, what we're looking at is more server-side stuff, but now we're going to have a look at client-side regex.
And a lot of it is used actually to sanitize HTML, which is, you know, kind of what it's actually supposed to do. It's good for finding substrings in HTML elements. So, if you want to find an element outside of it, regex is pretty good for that. And a lot of the sanitizers rely heavily on using correct regex. Now, you know, correct is something sometimes can actually backfire on us. And, you know, in different situations, in a perfect world, the regex will always work.
But especially when you're dealing with attacker input, where the attacker can craft these weird payloads with like embedded attributes and stuff, it can lead to some vulnerabilities. And I think this is a pretty cool write-up from this user here. All the write-ups will be in the description, by the way. I recommend you have a read of them yourself because I'm just kind of giving you a gloss over them. As you can see, they have this regex pattern here.
So, they created the regex pattern that's supposed to filter an attribute. Say this is like an onclick attribute or something. And basically, this is put straight into the regex. And what this will do is this will find the attribute name equals and then say like find these things in apostrophes, right? Obviously, there's a lot of issues with these kind of things, you know, like browsers will render stuff very differently, you know.
Sometimes you don't actually need the commas for the script to execute. Like if this was like a JavaScript URL, it would execute anyway. And, you know, the problem is also this greedy matching, which the person will go into. So, what is inside of this regex does is the dot, any character, the asterisk character, it will match basically from zero to an infinite amount of times, as many times as physically possible. And the question mark for, you remember, that's like the optional character, I call it like optional, but it actually does something very specific, which is it will make it match from like zero times to one time.
So, it will make it what's called lazy matching. There's two types of matching in regex, there's greedy matching and lazy matching. So, greedy is as many times as physically possible. Lazy matching is the least amount of times as possible. So, that basically is what switches it, and it's what what prevents the regex from consuming the comma after the question mark. That basically is because it'll stop consuming it as soon as possible.
Instead, if that wasn't there, it would consume the comma and just keep going forever, basically, because the dot asterisk is unbounded. And we have a look at something like this, for example, you know, they showed the regex working, you know, it goes down, but then, you know, you have these situations where the regex pattern, you know, it's not a perfect world anymore. The attacker can tailor different kind of things, and, you know, they showed here, for example, that this was inside of an actual tag itself, which means that the regex would eat this like comma, and then it would continue to the next one, which happened to be here.
So, it actually consume the entire tag. And this is because the regex didn't factor in that the actual attribute name could have been inside of another tag that the attacker controls. So, that is super interesting, and that's how that can lead to like the JavaScript vulnerability like this, for example. And, you know, it can also be, for example, inside of another tag, like so. So, it can be here. And then it can, you know, end up being consumed.
So, it's just kind of a difference of where the developer kind of thought the attribute would be versus where the attribute actually can be in the certain like rejects kind of manipulations you can do there. And yeah, there's different kind of things to see, but you know, this is a very cool example of how rejects can kind of go wrong, especially with XSS vulnerabilities. And I really recommend that you read this write-up.
And yeah, by the way, if you do enjoy my content, then you might want to have a look in the description at my community. It's a program that will give you access to my community where you can, you know, chat with me and other people in there, as well as get access to some of my labs, some of the interactive labs that I built specifically focused on uh server-side vulnerabilities and some client-side ones that I've personally modeled based by, you know, some of the research that I do basically every day.
And I think it's very cool product. So, if you do enjoy my videos, then there's a good chance you'll also enjoy it. So, check out the link in the description. Hope you enjoyed this video covering some of the write-ups around rejects. It's, you know, something that I'm working on as well now. I'm sharing it with you guys. I've been doing it for a few years now, and I think I have a bit of experience now. All the articles are going to be in the description, so go have a read of it.
And I also encourage you to practice rejects yourself. If you want to get good at identifying it, you know, writing simple rejects with the tool I showed you previously, it's called rejects101, will actually, I think, really start leveling up your skills in it and allow you to spot situations, you know, make you more conscious of situations where rejects has been built incorrectly, where we can deliver specifically crafted input that will bypass the reject that is being used.
[music] And that is a very important skill is, you know, being able to identify, you know, kind of like unit testing in programming. If you're not familiar, unit testing is like when you have a function, you will isolate it, and then you put a load of different inputs [music] into it to see how it kind of reacts. And that's kind of the same thing you want to do with regex. So, you want to take the bit of regex, you want to think about, you know, what kind of inputs can you put in here?
Almost like brainstorming. And just fire them out and see like, what does it do when this happens? What if, you know, it's inside the attribute? What if it's like, you know, these type of things? And I think that's a very cool practice that you can honestly do yourself as well. That's it for me. Thanks for watching, and happy hunting.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.