JavaScript Regular Expressions / lesson 5 of 11

Shorthand Unicode Properties

What IS a regular expression, and why should you care?

Play
Transcript

[00:00] So when we want to capture, say, uppercase A through Z and lowercase A through Z and 0 through 9, basically alphanumeric, that is our character class. Since we are dealing with some interesting Unicode characters here, I’m going to enable the U flag. When I save that, you can see it is, in fact, finding all of the alphanumeric values. Now, we can shorten that by just saying backslash W. So W is now our character class, and we escape it with the backslash. And when I save that, we’re going to get the exact same results.

[00:38] And if we want to negate that, we can create another character class, wrap that in there, and we just add our caret to negate. And now it’s going to select everything that’s not alphanumeric. And there’s actually a shorthand for this guy, which is just backslash capital W. So exact same results there. Now we can get digits with backslash D. So now we’ve got only the digits. And, of course, we can negate that in our own character class. Now it’s everything that’s not a digit. And then we can shorthand the negation with just a capital D. And there we go.

[01:12] Now we can use an S to get all of the white spaces. And we can go through the exact same pattern where we add our negation. And now we’re getting everything that’s not a white space. And, of course, there’s a shorthand for that, which is capital S. So that’s everything that’s not a white space. Those are, like, the most commonly used character class shorthands ever. But there is something really cool that we’re going to talk about. And that’s Unicode property classes, I guess is the word for it.

[01:43] Now this was introduced somewhere around 2020. But I don’t know if a lot of people use it. But we start off with a backslash P. And then we’re going to have our modifier in here. And we can do all sorts of stuff. So the first thing I’m going to do is I’m going to say symbol. And we’ll take a look at what’s happening. It’s capturing currency symbols, mathematic symbols, emojis. All of those are being captured. So that’s pretty cool. We can shorthand symbol with just an S. And, boom, we’ve got the exact same results.

[02:16] But then we can modify that further. So let’s say I only want currency symbols. You can do SC. Save that. And you can see it’s highlighting the dollar symbol and the euro symbol. We could do SM. And what that’s going to do is mathematical symbols. So you can see it’s getting the multiplier, the equals, the division, and so on. And there’s other modifiers we can add to this. I’ll definitely have a link in the description. Okay, so we can do punctuation. And now you can see it’s getting decimals or, well, maybe periods if it’s looking at them as a period. Parentheses, opening, closing quotes, and all that. That’s really cool. We can shorten that to just P and we get the exact same results.

[02:55] We can set that to PS, which is going to get the start or open punctuation. And in this case, it considers the open parentheses to be that. It does not consider the initial double quotes to be that. We’ll cover that in just a second. So that’s our PS. Then we can do PE for the end. And you can see it’s capturing the closing or the ending parentheses. So that’s cool. Now, for the double quotes, it considers that initial and final. So I’ve got an initial double quote there and a final double quote on each of those strings. So I say PI and I get the initial ones. I say PF for final and I get the final ones. And just to point out, we can take these and throw them into a new class. So I’m going to say I want PI and PF. Save that. And you can see I’m getting both of them now. So you can put these together just like any other character classes. It ends up making them really, really powerful.

[03:55] So another one we can do is letter. So give me all the letters. Cool. We can shorten that to just L. And then we can do something like I want all the lowercase letters with L, lowercase L. And there you go. We’ve got all the lowercase letters. And of course, I can do uppercase letters. So there we go. We’ve got all the uppercase letters. Another one we can do here is number. So very similar to digits. It’s capturing all the numbers. It’s also capturing some of these. I think those are considered number letters, the fractional elements there. So we can just shorten that to just N. We get the exact same results. And if I only wanted to get actually their number other is what they are, those fractional things. Sorry, I’m going to save that. And boom, I’ve only got those fractional symbols. Super cool.

[04:46] Now, we’ve got these foreign characters. They are Cyrillic and Greek. This guy is Cyrillic. And this guy is Greek. So if we want to get just those, we can use something called script. And now we’re actually going to give this property, we’re going to give it a value. I’m going to try to spell this right. Cyrillic. So I want everything that’s Cyrillic. And you can see it did, in fact, grab that. If I wanted just the Greek, we can get that. There’s just the Greek. And then we can also shorten this script to just SC. Oops, sorry. Lowercase SC. And there you go. We’ve got only the Greek letters. So that is a quick look at character shorthands. And the really, really powerful feature that is Unicode properties.

This lesson covers the most commonly-used character class shorthands and the powerful Unicode property escapes. Learn \w, \d, and \s (and their negated forms \W, \D, \S), then we’ll dive deep into Unicode property escapes with \p{...}. You’ll explore symbols (\p{Symbol}), punctuation (\p{Punctuation}, \p{PS}, \p{PE}, \p{PI}, \p{PF}), letters (\p{Letter}), numbers (\p{Number}), and even scripts (\p{Script=Cyrillic}, \p{Script=Greek}). These features let you match characters by their semantic meaning rather than by range.

app.js

app.ts
import output from "./output.js";
const str = `Aeiou (vowels)
$100 × 0.5% = €0.43
¼ ÷ ½ = ¾
😀 “Καλημέρα”
😴 “Спокойной”`;
let regex = /[a-zA-Z0-9]/g;
regex = /\w/g; // only alpha numeric
// regex = /[^\w]/gu; // no alpha numeric
// regex = /\W/gu; // no alpha numeric
// regex = /\d/gu; // only digits
// regex = /[^\d]/gu; // only digits
// regex = /\D/gu; // no digits
// regex = /\s/gu; // only whitespace
// regex = /[^\s]/gu; // only whitespace
// regex = /\S/gu; // no whitespace
// regex = /\p{Symbol}/gu; // only symbols
// regex = /\p{S}/gu; // only symbols
// regex = /\p{Sm}/gu; // only symbols math
// regex = /\p{Sc}/gu; // only symbols currency
// regex = /\p{Punctuation}/gu; // only puncutation
// regex = /\p{P}/gu; // only puncutation
// regex = /\p{Ps}/gu; // only puncutation start or open
// regex = /\p{Pe}/gu; // only puncutation end or close
// regex = /\p{Pi}/gu; // only initial puncutation
// regex = /\p{Pf}/gu; // only final puncutation
// regex = /[\p{Pi}\p{Pf}]/gu; // only initial and final puncutation
// regex = /\p{Letter}/gu; // only letters
// regex = /\p{L}/gu; // only letters
// regex = /\p{Lu}/gu; // only letters uppercase
// regex = /\p{Ll}/gu; // only letters lowercase
// regex = /\p{Number}/gu; // only numbers
// regex = /\p{N}/gu; // only letters
// regex = /\p{No}/gu; // only letters
// regex = /\p{Script=Cyrillic}/gu;
// regex = /\p{Script=Greek}/gu;
// regex = /\p{sc=Greek}/gu;
output(str, regex);

Share this post on:

Previous
Character Classes
Next
Capturing Groups