Add u8"..." and u8R"(...)" cases to test_locations in test.cpp, guarded
by #if __cplusplus >= 201703L. Add an options file forcing
-std=c++17 so the u8 char8_t-era rows are actually extracted.
Plain "..." rows are unchanged. New u8/u8R rows in locations.expected
show correct columns computed by regexpContentOffset: offset 3 for
`u8"` and offset 5 for `u8R"(`. No change to regexpContentOffset was
needed.
Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
POSIX bracket expressions ([:name:], [.x.], [=x=]) are only valid nested
inside another character class in std::regex. The three predicates now
additionally require that at position `start` we are inside an outer
character class, approximated by requiring more non-escaped `[` than
non-escaped `]` before `start`. A well-formed POSIX bracket contributes
one `[` and one `]` at/after `start`, so the check is unaffected by
earlier POSIX brackets.
Also stop emitting individual charSetTokens for characters inside a
POSIX bracket (`inAnyPosixBracket` guard), which removes the spurious
`[RegExpCharacterRange] :]-z` on [[:alpha:]-z].
Post-fix, unnested `[:digit:]`, `[:alpha:]`, and `a[🅱️]c` are correctly
parsed as ordinary character classes / literal sequences rather than
`RegExpNamedCharacterProperty`; and the [[:alpha:]-z] range no longer
crosses the POSIX bracket boundary.
Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
Add corpus cases: bare [:alpha:], mid-pattern a[🅱️]c, POSIX class as
range endpoint [[:alpha:]-z], malformed [[:alpha], leading literal ]
combined with a POSIX class []a[:alpha:]], three POSIX classes in one
class, additional names [[:xdigit:]] / [[:blank:]] / [[:cntrl:]] /
[[:graph:]], and the integration case combining all of the above.
The regenerated parse.expected reveals genuine mis-parses that this
commit deliberately records (fixes in Commit 6):
- unnested [:alpha:] and [🅱️] are treated as RegExpNamedCharacterProperty
even though POSIX brackets are only valid inside a character class,
- [[:alpha:]-z] emits a spurious [RegExpCharacterRange] ":]-z" instead
of a POSIX class + literal '-' + 'z'.
Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
In ECMAScript std::regex, \0 matches NUL. The previous escapedCharacter
arms all rejected it: the final arm's `not exists(getChar(start+1).toInt())`
guard fails for "0", and the numbered-backref arm excludes 0. Add an
explicit \0 case (end=start+2) guarded so the following character is not
a digit, mirroring EcmaRegExp.escapedCharacter in PR #22200. Add corpus
case "a\\0b" to test.cpp; \0 now parses as RegExpEscape.
Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
std::regex ECMAScript mode does not support \p{Name}; \p tokenizes as an
identity escape. Remove pStyleNamedCharacterProperty and its call sites;
namedCharacterProperty now covers only POSIX brackets ([:name:], [.x.],
[=x=]) and namedCharacterPropertyIsInverted keeps only the [[:^name:]]
case (fixing the offset from start+3 to start+2 for POSIX). Update
RegExpNamedCharacterProperty getName/isInverted docs to describe POSIX
brackets. Remove the four \p{...}/\P{...} corpus lines from test.cpp;
keep a plain [a-f\d]+ case for the class-with-escape shape.
Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).