are an important concept in formal language theory. They are a way to describe a possibly infinite set of character strings (called a language). A regular expression, at its core, needs the following features:
A set of characters that can be used in the language, called the alphabet.
Concatenation: ab means "the character a followed by the character b".
Union: a|b means "either a or b".
Kleene star: a* means "zero or more a characters".
Assuming a finite alphabet (such as the 26 letters of the English alphabet, or the entire Unicode character set), all regular languages can be generated by the features above. Of course, many patterns are very tedious to express this way (such as "10 digits" or "a character that's not a space"), so JavaScript regular expressions include many shorthands, introduced below.
Note: JavaScript regular expressions are in fact not regular, due to the existence of
(regular expressions must have finite states). However, they are still a very useful feature.
A regular expression is typically created as a literal by enclosing a pattern in forward slashes (/):
js
const regex1 = /ab+c/g; Regular expressions can also be created with the
constructor:
js
const regex2 = new RegExp("ab+c", "g"); They have no runtime differences, although they may have implications on performance, static analyzability, and authoring ergonomic issues with escaping characters. For more information, see the
reference.
Flags are special parameters that can change the way a regular expression is interpreted or the way it interacts with the input text. Each flag corresponds to one accessor property on the RegExp object.
FlagDescriptionCorresponding propertydGenerate indices for substring matches.
gGlobal search.
iCase-insensitive search.
mMakes ^ and $ match the start and end of each line instead of those of the entire string.
sAllows . to match newline characters.
u"Unicode"; treat a pattern as a sequence of Unicode code points.
vAn upgrade to the u mode with more Unicode features.
yPerform a "sticky" search that matches starting at the current position in the target string.
The i, m, and s flags can be enabled or disabled for specific parts of a regex using the
syntax.
The sections below list all available regex syntaxes, grouped by their syntactic nature.
Assertions are constructs that test whether the string meets a certain condition at the specified position, but not consume characters. Assertions cannot be
.
Buffer boundary assertion: \A, \z, \Z
Asserts that the current position in the string is strictly at the start or end of the entire string (\Z also allows a trailing newline), regardless of the presence of the m flag.
Input boundary assertion: ^, $
Asserts that the current position is the start or end of input, or start or end of a line if the m flag is set.
Lookahead assertion: (?=...), (?!...)
Asserts that the current position is followed or not followed by a certain pattern.
Lookbehind assertion: (?<=...), (?<!...)
Asserts that the current position is preceded or not preceded by a certain pattern.
Word boundary assertion: \b, \B
Asserts that the current position is a word boundary.
Atoms are the most basic units of a regular expression. Each atom consumes one or more characters in the string, and either fails the match or allows the pattern to continue matching with the next atom.
Matches a previously matched subpattern captured with a capturing group.
Matches a subpattern and remembers information about the match.
Character class: [...], [^...]
Matches any character in or not in a set of characters. When the
flag is enabled, it can also be used to match finite-length strings.
Character class escape: \d, \D, \w, \W, \s, \S
Matches any character in or not in a predefined set of characters.
Matches a character that may not be able to be conveniently represented in its literal form.
Matches a specific character.
Overrides flag settings in a specific part of a regular expression.
Matches a previously matched subpattern captured with a named capturing group.
Named capturing group: (?<name>...)
Matches a subpattern and remembers information about the match. The group can later be identified by a custom name instead of by its index in the pattern.
Matches a subpattern without remembering information about the match.
Unicode character class escape: \p{...}, \P{...}
Matches a set of characters specified by a Unicode property. When the
flag is enabled, it can also be used to match finite-length strings.
Matches any character except line terminators, unless the s flag is set.
These features do not specify any pattern themselves, but are used to compose patterns.
Matches any of a set of alternatives separated by the | character.
Quantifier: *, +, ?, {n}, {n,}, {n,m}
Matches an atom a certain number of times.
Escape sequences in regexes refer to any kind of syntax formed by \ followed by one or more characters. They may serve very different purposes depending on what follow \. Below is a list of all valid "escape sequences":
Escape sequenceFollowed byMeaning\ANone
\BNone
\DNone
representing non-digit characters\P{, a Unicode property and/or value, then }
Unicode character class escape
representing characters without the specified Unicode property\SNone
representing non-white-space characters\WNone
representing non-word characters\ZNone
\bNone
; inside
, represents U+0008 (BACKSPACE)\cA letter from A to Z or a to zA
representing the control character with value equal to the letter's character value modulo 32\dNone
representing digit characters (0 to 9)\fNone
representing U+000C (FORM FEED)\k<, an identifier, then >A
\nNone
representing U+000A (LINE FEED)\p{, a Unicode property and/or value, then }
Unicode character class escape
representing characters with the specified Unicode property\q{, a string, then a }Only valid inside
; represents the string to be matched literally\rNone
representing U+000D (CARRIAGE RETURN)\sNone
representing whitespace characters\tNone
representing U+0009 (CHARACTER TABULATION)\u4 hexadecimal digits; or {, 1 to 6 hexadecimal digits, then }
representing the character with the given code point\vNone
representing U+000B (LINE TABULATION)\wNone
representing word characters (A to Z, a to z, 0 to 9, _)\x2 hexadecimal digits
representing the character with the given value\zNone
\0None
representing U+0000 (NULL)\ followed by 0 and another digit becomes a
, which is forbidden in
. \ followed by any other digit sequence becomes a
.
In addition, \ can be followed by some non-letter-or-digit characters, in which case the escape sequence is always a
representing the escaped character itself:
\$, \(, \), \*, \+, \., \/, \?, \[, \\, \], \^, \{, \|, \}: valid everywhere
\-: only valid inside
\!, \#, \%, \&, \,, \:, \;, \<, \=, \>, \@, \`, \~: only valid inside
The other
characters, namely space character, ", ', _, and any letter character not mentioned above, are not valid escape sequences. In
, escape sequences that are not one of the above become identity escapes: they represent the character that follows the backslash. For example, \a represents the character a. This behavior limits the ability to introduce new escape sequences without causing backward compatibility issues, and is therefore forbidden in Unicode-aware mode.
Specification
ECMAScript® 2027 Language Specification# prod-PatternCharacter
ECMAScript® 2027 Language Specification# prod-CharacterClass
ECMAScript® 2027 Language Specification# prod-RegularExpressionModifiers
ECMAScript® 2027 Language Specification# prod-Disjunction
ECMAScript® 2027 Language Specification# prod-CharacterClassEscape
ECMAScript® 2027 Language Specification# prod-Atom
ECMAScript® 2027 Language Specification# prod-RegExpUnicodeEscapeSequence
ECMAScript® 2027 Language Specification# prod-CharacterEscape
ECMAScript® 2027 Language Specification# prod-DecimalEscape
ECMAScript® 2027 Language Specification# prod-AtomEscape
ECMAScript® 2027 Language Specification# prod-Assertion
ECMAScript® 2027 Language Specification# prod-Quantifier
ECMAScript® 2027 Language Specification# sec-patterns-static-semantics-early-errors
Regular Expression Buffer Boundaries for ECMAScript# sec-patterns
guide