Commit Graph

9 Commits

Author SHA1 Message Date
copilot-swe-agent[bot]
8d53e365d6 Commit 5: Port missing POSIX-bracket parse cases from PR #22200
Add corpus cases: bare [:alpha:], mid-pattern a[🅱️]c, POSIX class as
range endpoint [[:alpha:]-z], malformed [[:alpha], leading literal ]
combined with a POSIX class []a[:alpha:]], three POSIX classes in one
class, additional names [[:xdigit:]] / [[:blank:]] / [[:cntrl:]] /
[[:graph:]], and the integration case combining all of the above.

The regenerated parse.expected reveals genuine mis-parses that this
commit deliberately records (fixes in Commit 6):
- unnested [:alpha:] and [🅱️] are treated as RegExpNamedCharacterProperty
  even though POSIX brackets are only valid inside a character class,
- [[:alpha:]-z] emits a spurious [RegExpCharacterRange] ":]-z" instead
  of a POSIX class + literal '-' + 'z'.

Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
2026-07-21 12:21:45 +00:00
copilot-swe-agent[bot]
76e3118aba Commit 2: Fix \0 escape parsing
In ECMAScript std::regex, \0 matches NUL. The previous escapedCharacter
arms all rejected it: the final arm's `not exists(getChar(start+1).toInt())`
guard fails for "0", and the numbered-backref arm excludes 0. Add an
explicit \0 case (end=start+2) guarded so the following character is not
a digit, mirroring EcmaRegExp.escapedCharacter in PR #22200. Add corpus
case "a\\0b" to test.cpp; \0 now parses as RegExpEscape.

Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
2026-07-21 12:18:23 +00:00
copilot-swe-agent[bot]
448ae6af53 Commit 1: Remove \p{...}/\P{...} named-character-property handling
std::regex ECMAScript mode does not support \p{Name}; \p tokenizes as an
identity escape. Remove pStyleNamedCharacterProperty and its call sites;
namedCharacterProperty now covers only POSIX brackets ([:name:], [.x.],
[=x=]) and namedCharacterPropertyIsInverted keeps only the [[:^name:]]
case (fixing the offset from start+3 to start+2 for POSIX). Update
RegExpNamedCharacterProperty getName/isInverted docs to describe POSIX
brackets. Remove the four \p{...}/\P{...} corpus lines from test.cpp;
keep a plain [a-f\d]+ case for the class-with-escape shape.

Bundle: github/codeql-action codeql-bundle-v2.26.1 (CodeQL CLI 2.26.1).
2026-07-21 12:16:37 +00:00
copilot-swe-agent[bot]
bfa65301f1 Commit 9: Fix hasLocationInfo for all string literal forms; fix unicode escape
RegexTreeView.qll - hasLocationInfo fix (RULE 4):
- Add regexpContentOffset(RegExp re) private helper that computes the correct
  number of source chars before the first content char, from getValueText():
  - Plain "...":          offset 1
  - L"...":               offset 2 (L" prefix)
  - u8"...":              offset 3 (u8" prefix)
  - R"(...)":             offset 3 (R"( opener)
  - R"x(...)x":          offset 4 (R"x( with custom delim)
  - LR"(...)":            offset 4 (LR"()
  - Uses vt.matches("%R\"%(%") to detect raw strings; finds '(' position
  - For non-raw, finds '"' position
- Replace hardcoded `+ 1` with `+ regexpContentOffset(re)` in hasLocationInfo
- Update comment: documents plain/encoding/raw/combined forms, notes that
  escaped non-raw strings are approximate (mirroring Java/Python)

test.cpp:
- Fix r_uni: change "\\u{9879}" to "\\u9879" — braced \u{...} is not standard
  C++ regex syntax; the 4-digit \uHHHH form is valid ECMAScript \u escape

Regenerated all 3 .expected files via codeql test run --learn (CodeQL 2.26.1).
All 3 tests pass.

Location test confirms:
- Plain "a+b" (line 144): litStartCol 22 → termStartCol 23 (22+1)  UNCHANGED ✓
- Raw R"(a+b)" (line 147): litStartCol 20 → termStartCol 23 (20+3)  FIXED ✓
- Raw R"x(a+b)x" (line 156): litStartCol 21 → termStartCol 25 (21+4) FIXED ✓
- L"a+b" (line 159): litStartCol 22 → termStartCol 24 (22+2)         FIXED ✓
- LR"(a+b)" (line 162): litStartCol 26 → termStartCol 30 (26+4)      FIXED ✓
2026-07-21 10:36:49 +00:00
copilot-swe-agent[bot]
e7ccb1f2ce Commit 8: Add location tests exposing term-location off-by-one (pre-fix)
Add location test cases to test.cpp covering various C++ string literal forms:
- r_plain:     "a+b"          — plain, offset 1 (already correct)
- r_raw:       R"(a+b)"       — raw, offset 3 (R"( = 3); currently uses 1 (WRONG)
- r_raw2:      R"(\s+$)"      — raw with metacharacters; currently wrong
- r_raw3:      R"(\(([,\w]+)+\)$)" — complex raw; currently wrong
- r_raw4:      R"x(a+b)x"    — custom-delimiter raw, offset 4; currently uses 1 (WRONG)
- r_wide:      L"a+b"         — L" prefix, offset 2; currently uses 1 (WRONG)
- r_wide_raw:  LR"(a+b)"      — combined LR"( prefix, offset 4; currently uses 1 (WRONG)
- r_esc1:      "\\s+"         — escape-containing plain; offset 1 (correct)
- r_esc2:      "a\\.b"        — escaped dot; offset 1 (correct)

Added std::wregex typedef to stubs (for L"..." and LR"(...)" wide-char literals).

New locations.ql query reports per-term: litStartCol, valueStart, valueEnd,
termStartCol, termEndCol — restricted to test_locations() function.

Generated locations.expected via codeql test run --learn (CodeQL 2.26.1).
All 3 tests pass. The expected output captures the CURRENT (pre-fix) columns:
raw/prefixed rows show termStartCol = litStartCol + 1 (wrong);
plain rows show litStartCol + 1 (correct).
Commit 9 changes the raw/prefixed rows to their correct values.
2026-07-21 10:27:55 +00:00
copilot-swe-agent[bot]
9fd7e949c8 Commit 6: Add POSIX-bracket extension tests (pre-fix, capturing incorrect tokenization)
Add new test cases to test.cpp exercising POSIX bracket expressions inside
character classes as supported by std::regex ECMAScript mode:
- Single POSIX classes: [[:alpha:]], [[:digit:]], [[:space:]], [[:upper:]],
  [[:lower:]], [[:alnum:]], [[:print:]], [[:punct:]]
- Negated outer bracket: [^[:space:]]
- Mixed: regular char + POSIX class [a[:space:]]
- POSIX collating symbol [[.a.]]
- POSIX equivalence class [[=a=]]
- Mixed POSIX + range [[:alpha:]0-9]

The existing r_posix1-r_posix4 cases remain untouched. New cases use distinct
names (r_posix_alpha, r_posix_digit, etc.) to avoid redeclaration.

Generated .expected via codeql test run --learn (CodeQL 2.26.1); all 2 tests
pass. The expected output captures the CURRENT (pre-fix) tokenization — commit 7
will change these rows when correct POSIX nesting is implemented.
2026-07-21 10:17:14 +00:00
copilot-swe-agent[bot]
67ff5896a8 Commit 5: Remove Ruby-only parser features; shift to ECMAScript dialect
- ParseRegExp.qll:
  - specialCharacter: remove \A, \Z, \z, \G (Ruby-only anchors); keep only \b, \B
  - firstPart: remove \A arm; keep only ^ for start-of-string
  - lastPart: remove \Z, \z arms; keep only $ for end-of-string
  - nonCapturingGroupStart: remove < from char set (handled by namedGroupStart
    and lookbehindAssertionStart; was causing mis-tokenization of lookbehind)
  - namedGroupStart: remove Ruby-only single-quote (?'name'...) branch; keep
    only ECMAScript (?<name>...) syntax
- RegexTreeView.qll:
  - RegExpCharacterClassEscape: remove h, H (Ruby hex-digit classes); keep
    d, D, s, S, w, W (ECMAScript set)
  - RegExpAnchor: change to [^, $] only (remove \A, \Z, \z)
  - RegExpDollar: change to $ only (remove \Z, \z)
  - RegExpCaret: change to ^ only (remove \A)
- test.cpp:
  - r_cc3: replace \A[+-]?\d+ with ^[+-]?\d+ (ECMAScript caret)
  - r_meta5 \h\H: removed (Ruby-only hex classes)
  - r_anc1 \Gabc: removed (\G not in ECMAScript)
- Regenerated parse.expected and regexp.expected via codeql test run --learn
  (CodeQL 2.26.1); all 2 tests pass.
2026-07-21 10:15:58 +00:00
copilot-swe-agent[bot]
96c7e2dbc5 Commit 4: Curate corpus for C++ ECMAScript; regenerate .expected
Content-only edits to test.cpp to remove Ruby-only syntax that has no
C++ ECMAScript equivalent:

1. Removed: "a{,8}" — Ruby-only "{,n}" no-lower-bound quantifier
   (ECMAScript requires an explicit lower bound in {n,m})

2. Removed: second ".*" for /.*/m — mode flags in C++ are construction-site
   arguments (e.g. std::regex::multiline), not part of the pattern string.
   This mode variant was the exact same pattern string as r_meta1 so it
   added no coverage.

3. Removed: "(?'foo'fo+)" — Ruby single-quote named-group form.
   ECMAScript only supports the angle-bracket form (?<name>...).

Left in place for now (removed with the parser in commit 5):
  \\A, \\z, \\G, \\h, \\H — Ruby-only anchor/escape classes.
  The POSIX-bracket cases are kept through commit 7.

.expected regenerated by:
  codeql test run --learn --search-path=. cpp/ql/test/library-tests/regex/
  (CodeQL CLI 2.26.1)
All 2 tests passed.
2026-07-21 10:12:21 +00:00
copilot-swe-agent[bot]
d7a8f48c1c Commit 3: Add test.cpp corpus wrapped in std::regex stubs; generate .expected
Introduces cpp/ql/test/library-tests/regex/test.cpp with the Ruby regex
corpus mechanically translated to C++ string literals and wrapped in
std::regex construction calls using fully self-contained in-file stubs
(no #include of any external/standard header).

Stub surface:
  std::basic_regex<CharT>  constructor(const char*), constructor(const char*, int),
                           assign(const char*)
  std::regex               typedef for basic_regex<char>
  std::regex_match / regex_search / regex_replace  free functions

1:1 corpus mapping from regexp.rb:
  - All Ruby /regex/ literals translated to C++ "string" literals
    (each backslash doubled: /\d/ -> "\\d")
  - Dropped: /#{A}bc/ (Ruby string interpolation, no string-literal form)
  - Kept: a{,8}, .*m, \A, \z, \G, \h\H, (?'foo'...) — these are removed
    in commits 4 and 5 (keeping them here preserves 1:1 mapping)

Also removes `abstract` from RegExp class in ParseRegExp.qll so that all
StringLiterals are regex candidates (trivial syntactic gate; no dataflow).

.expected generated by:
  codeql test run --learn --search-path=. cpp/ql/test/library-tests/regex/
  (CodeQL CLI 2.26.1, bundled C/C++ extractor)
All 2 tests passed.
2026-07-21 10:11:27 +00:00