JavaScript Regular Expressions / lesson 8 of 11

Backreferences

Backreferences let you match a second occurrence of text captured by an earlier group.

Play
Transcript

[00:00] When we wanna find a second instance of a capture group, we can utilize something called back references. So to illustrate that, I’m gonna say it was the, the thing. So I’ve got the there twice, it’s a very common typo. Now in a regular expression, first thing we’re gonna do is we’re gonna capture the. So we see that on the right, we’ve got both from there. I’m gonna say it’s possibly followed by a white space. Now we can see we’ve got the white space there. Now to identify the second instance of THE, I’m gonna use a back reference. And what we do is we do a backslash and then the index of the capture group. So if you remember our matches, zero is gonna be the full input. So this is gonna be capture group one, that is its index. So when I save that, now we’ve got both of them and that’s great.

[00:54] But if I wanna separate them, what I can do is use a look ahead. So I say look ahead here and now I wanna capture the thing on the left. I wanna capture the thing that is repeated. That’s the first THE if it’s followed by another THE. So we could actually clean this guy up if we find that instance by saying string dot replace rejects and just replace it with an empty string. And if I load that up in the console, we can see that we have replaced it with it was the thing. Now let’s say I did it was the thing thing. Well, our current workaround is not gonna do it. So let’s just say it’s a word of one or more characters. So now we’ve got the and thing captured. And when we do our exact same string replace, we get it was the thing. So that works great. That’s a back reference.

[01:50] A really common use case for back references is working with HTML. So let’s say I’ve got a bold tag and I close the bold tag. And then here I just say bold text, right? That’s not gonna work great in our little output guy here. So we’re gonna stick with the console. We’re just gonna update this guy a bit. So now our word, whatever it is, is gonna be wrapped in less than and greater than. So we’re trying to find that HTML tag. We won’t need the look ahead here. So we’ll just say less than and we’ll escape the forward slash. And we’ll put the greater than there. Save that and it’s looking good so far. But here in the middle, we have to say, you know, one or more characters. So we’ll do that. So it’s actually, you can see it has selected that text in the UI. But now what we want to do is get our second capture group. So we’re replacing everything with our second capture group. This was our first capture group, which we’re back referencing here. So this part in the middle is our second capture group. So when I save that, I get the text inside of the tags. So I’ve stripped out the HTML.

[03:09] So just like with other capture groups, we can actually name our back reference groups. So it’s a slightly different implementation or syntax for the back reference. But the first one is actually the exact same. So I just say question mark. I’m going to call this tag. And then instead of this backslash one, I can just reference the name of it. And to identify it as a named group in my back reference, I precede that with a backslash K. And now what I could do is say const match equals rejects dot execute string. And then just console log. And we’ll say tag. We’ll get that a little arrow. And we’ll say match dot groups dot tag. When I save that, we get our bold tag. So if this was block quote, we get our block quote. So cool. That is a quick look at how to use back references and how to use back references with named groups, which is, again, another really powerful thing that you can take advantage of.

Backreferences let you match a second occurrence of text captured by an earlier group. This lesson covers the syntax \1, \2, etc., and shows practical use cases: finding duplicate words (like “the the”), stripping HTML tags by matching opening and closing tags together, and removing duplicate patterns from text. You’ll also learn named backreferences and how they integrate with named capturing groups for cleaner, more maintainable regex.

app.js

app.ts
import output from "./output.js";
// ─── Backreferences ──────────────────────────────────
let str = "it was the the thing";
let regex = /(the)\s?/g;
regex = /(the)\s?\1/g;
regex = /(the)\s?(?=\1)/g;
console.log(str.replace(regex, ""));
str = "it was the the thing thing";
regex = /(\w+)\s?(?=\1)/g;
console.log(str.replace(regex, ""));
str = `<b>Bold text</b>`;
regex = /<(\w+)>(.*)<\/\1>/g;
console.log(str.replace(regex, "$2"));
str = `<b>Bold text</b>`;
regex = /<(?<tag>\w+)>(.*?)<\/\k<tag>/g;
match = regex.exec(str);
console.log("tag ->", match.groups.tag);
output(str, regex);

Share this post on:

Previous
Lookaheads and LookBehinds
Next
Word Boundaries