* docs/tests: fix constructs index regression from #1196; add ref page to sidebar; extend & ref tests
- Restore extension-less links and the text_blocks entry in
docs/lpc/constructs/index.md (the PR was recreated from a pre-Docusaurus
branch and reintroduced .html links, which fail the docs build under
onBrokenLinks: 'throw', and dropped text_blocks)
- Add lpc/constructs/ref to the hand-authored sidebar in docs/sidebars.ts
- Drop the stale VitePress 'layout: doc' frontmatter from ref.md
- Extend testsuite/single/tests/operators/ref.lpc: & in parameter
declarations, ref keyword in foreach, and bitwise &/&= non-regression
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Skf9CfHwAngWorDNDzPKWp
* vm: fix string foreach crash and off-by-one ref reads; support += / -= on string chars
Three related defects around the shared string-codepoint lvalue
(global_lvalue_codepoint), all reproduced on the unfixed binary:
- Nested foreach over strings SEGFAULTED: the iteration cursor lived in
the shared global, so the inner loop's F_EXIT_FOREACH reset the
iterator out from under the outer loop (null deref in
post_index_to_offset). The EGC cursor now lives in each loop's own
stack slot (the T_NUMBER slot under the loop variable); the ref case
re-arms the shared codepoint lvalue every iteration, and exit only
clears the global when it still points at this loop's string slot.
- foreach (int ref c in str) read the WRONG characters: the shared
index was advanced before the body ran, so every read through the ref
was off by one ('abc' summed to 197 instead of 294, and the final
iteration read one past the end). The global index now stays on the
current character for the whole body. Writes through the ref still go
to the loop's by-value stack copy and never reach the iterated
variable -- semantics pinned by tests/operators/foreach.lpc.
- s[i] += n / s[i] -= n threw "Bad Argument 1 to +=()": F_ADD_EQ and
f_sub_eq handled buffer byte lvalues (T_LVALUE_BYTE) but not string
codepoint lvalues (T_LVALUE_CODEPOINT), even though ++/--/= worked.
Both now route through a new codepoint_lvalue_add() helper and produce
the resulting character as the rvalue.
Regression tests: nested (plain / ref / multi-byte UTF-8) string foreach
and ref-read correctness in tests/operators/foreach.lpc; compound
assignment on string chars (incl. reverse index, rvalue result, and
non-number rhs error) in tests/operators/string_index.lpc. Verified on
the unfixed binary: the nested-foreach test segfaults, the others fail.
Full LPC suite passes 2x on clang ASan/UBSan Debug and 2x on
RelWithDebInfo.
Docs: correct the ref.md note on foreach-over-strings; document the
string-char lvalue rules in AGENTS.md (section 8 + audit checklist 9)
and the ref/& feature in README.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Skf9CfHwAngWorDNDzPKWp
* vm: foreach ref over strings shares the s[i] char-lvalue logic; thorough ref tests
Arming a string-char lvalue now goes through one shared helper,
aim_lvalue_codepoint(): it validates the target EGC (out-of-bounds and
multi-codepoint error cleanly, same messages as s[i]) and points the
shared codepoint state at it. Both s[i] lvalues (push_indexed_lvalue)
and foreach ref loop variables use it, and every write consumer already
funnels into assign_lvalue_codepoint() -- so ref loop chars follow
exactly the s[i] rules:
- single-codepoint characters (and EGCs up to 4 bytes, which index as
their first codepoint) keep working: reads deliver the character,
assignments through the loop variable succeed
- wider EGCs (flag emoji, ZWJ sequences) now raise the catchable
"Indexed character is multi-codepoint" error when the ref loop
reaches them, instead of silently reading as -1; the non-ref form
still iterates and delivers -1 for such clusters
tests/operators/ref.lpc is now a thorough pass-by-reference suite:
ref/& parameter declarations and call arguments, by-value contrast,
ref forwarding through call chains, call-site refs to array elements /
mapping values / string chars (forward and reverse index, write-back),
foreach ref over arrays / mapping values / strings (read correctness,
multi-byte codepoints, by-value write semantics, assignment error
parity, the multi-codepoint error, non-ref contrast), compile-time
rejections via generated sources (ref outside an argument list, ref to
a range), and bitwise &/&= non-regression -- 39 checks.
Full LPC suite passes 2x on RelWithDebInfo and 2x on a clean clang
ASan/UBSan Debug build.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Skf9CfHwAngWorDNDzPKWp
* tests: cover ref across all LPC types
Extend tests/operators/ref.lpc to 58 checks: ref parameters of every
value type (int, float, string, array, mapping, object, function,
buffer, class, mixed) verifying reassignment propagates; the by-value
contrast for reference-typed containers (member writes propagate,
reassignment doesn't); call-site refs to buffer bytes (forward and
reverse index) and class members (both . and -> spellings); foreach
ref over mapping values of mixed types; the compile-time rejection of
a ref mapping KEY in foreach; and the clean runtime error for foreach
over a non-iterable type (buffer).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Skf9CfHwAngWorDNDzPKWp
* buffers behave like byte arrays: foreach, strict 0..255 bytes, string/array promotion, to_buffer()
Buffers are now first-class byte containers:
- foreach iterates a buffer like an array, delivering each byte as an
unsigned int 0..255. A ref loop variable mutates the buffer in place;
each ref carries its OWN T_LVALUE_BYTE (in ref->sv), so nested buffer
ref loops and b[i] lvalues in the body can't alias each other.
T_LVALUE_BYTE consumers now read the lvalue's own pointer/subtype
instead of reaching for the shared global_lvalue_byte (which remains
only as the scratch instance b[i] arms).
- every LPC byte write path (=, ++, --, +=, -=) range-checks the result:
a value outside 0..255 raises "Buffer byte value out of range" and
leaves the byte unchanged, instead of silently truncating/wrapping.
+= / -= on bytes also yield the resulting value as their rvalue.
- strings and arrays of ints 0..255 PROMOTE to buffers: a new
to_buffer() efun (registered like to_int/to_float) is wrapped around
the rhs by do_promotions() / rule_expr_assign for 'buffer b = str',
'b += str', initializers, and 'b + str'; range assignment and the
runtime + / += paths convert unpromoted (mixed) values through the
same svalue_to_buffer_bytes() helper. A string contributes its raw
UTF-8 bytes; an array must hold only ints 0..255 (validated before
allocation) or the conversion errors with the target unchanged.
- fixed a pre-existing ref_t leak: a foreach ref loop variable reused by
a re-entered inner loop leaked one ref per outer iteration (flagged by
the debug memory checker as 'Found temporary block: make_ref').
tests/operators/buffer_bytes.lpc pins the byte-range semantics, + / +=
concatenation, all promotion forms, and the error paths (106 checks);
foreach.lpc and ref.lpc pin buffer iteration and ref-loop independence;
buffer_range_assign.lpc gains range-read pins. New docs for to_buffer
(sidebar regenerated) and a rewritten lpc/types/buffer page.
Full LPC suite passes 2x on RelWithDebInfo and 2x on a clean clang
ASan/UBSan Debug build (no ref-checker warnings); GTest suite 312/312.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Skf9CfHwAngWorDNDzPKWp
---------
Co-authored-by: Claude <noreply@anthropic.com>
Add a 'UTF-8 Native Strings' section to lpc/types/strings.md covering
what the driver actually implements: lengths and positions are measured
in extended grapheme clusters (UAX #29), indexing yields code points and
errors on multi-code-point clusters, ranges/explode/strsrch operate on
character boundaries, display width (strwidth, UAX #11) vs length,
\uXXXX and surrogate-pair escapes, UTF-8 validity requirements, and the
encoding boundary (set_encoding for connections, string_encode /
string_decode / buffer_transcode elsewhere). Note in the old
sub-ranging section that positions are characters, not bytes.
Clarify sizeof() (string = grapheme clusters, buffer = bytes) and
strsrch() (character offsets, character-boundary matches).
Every documented example is pinned by a new testsuite file,
testsuite/single/tests/compiler/utf8_doc_examples.lpc, verified against
a freshly built driver (19 checks). Notably replace_string() is
byte-oriented, so it is deliberately NOT listed among the
grapheme-aware operations.
Claude-Session: https://claude.ai/code/session_01TSzcESzU9947zkGzQ6SMmE
Co-authored-by: Claude <noreply@anthropic.com>
* docs: add local full-text search and a contributor README
Add @easyops-cn/docusaurus-search-local to the Docusaurus site so the
docs get an offline search bar (index built at build time, no external
service). English and zh-CN pages are both indexed, and matched terms
are highlighted on the target page.
Add docs/README.md describing the Docusaurus setup, local dev/build
commands, search behavior, directory layout, and gotchas; exclude it
from the published site alongside CLAUDE.md. Point the root README's
docs/ entry at it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TSzcESzU9947zkGzQ6SMmE
* docs: remove dead framework leftovers, fix index generation, complete the nav
Delete the VitePress (.vitepress/) and Jekyll (_layouts/, css/) leftovers,
the one-shot migration scripts (fix_md_header.py, fix_seealso.py), and the
stale keywords.json snapshot; prune the matching .gitignore entries and
docusaurus exclude patterns.
Rewrite gen_index.py for Docusaurus: it emitted dead .html links and
legacy 'layout: doc' frontmatter, choked on non-markdown entries, and
dropped nested categories — regenerating an index would have broken it.
It now emits the extension-less links the site actually uses, links
nested category indexes (restoring apply/* on the zh-CN index), and
refuses to run on the docs root. Fix update_index.sh's copy-paste titles
(zh-CN efun/build were titled 'APPLY'), stop it clobbering the
hand-written lpc/index.md, and cover cli/. Regenerated indexes pick up
the missing driver/ffi-plan entry. add_missing_efuns.py now takes the
keywords.json path as an argument instead of requiring a stale copy.
Move CNAME and the Google site-verification file into static/ so they
actually reach the published build output.
Complete the sidebar: link the CLI category to cli/index and add the
missing portbind/symbol/generate_keywords pages, and expose the
previously orphaned stdlib section under Reference.
Promote onBrokenLinks to 'throw' now the build is warning-free, and drop
the empty Demo section from the landing page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TSzcESzU9947zkGzQ6SMmE
* docs: strip legacy 'layout: doc' frontmatter from all pages
Mechanical sweep removing the Jekyll-era 'layout: doc' line from every
doc page's frontmatter (Docusaurus ignores it), and the matching line
from the templates in docs/CLAUDE.md so new pages don't reintroduce it.
No content changes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TSzcESzU9947zkGzQ6SMmE
---------
Co-authored-by: Claude <noreply@anthropic.com>
Neither master's hand-written scanner nor the flex lexer ever accepted
exponents ("2.5e2" lexed as REAL(2.5) IDENT(e2)) -- but the hand-written
EBNF already documented the form and it should exist. Two lexer rules
add it: an optional exponent on the fraction form, and a bare-exponent
form ("1e3"); "1e" without exponent digits stays NUMBER(1) IDENT(e),
"1..5" stays a range, hex "0x1E3" is untouched, and '_' separators work
in all parts. strtod() already converted exponents, so only the
patterns changed.
Pinned at every layer: syntax_literals.lpc (2.5e2, 2.5e-2, 1.5E+2,
1e3, 1_2e4), the JS tokenizer + generated tmLanguage (with node tests
for the exponent forms and the "1e" non-exponent), the lexical EBNF
(realLiteral | exponent rules, replacing a half-aspirational entry),
and docs/lpc/types/float.md (which also wrongly claimed single
precision -- LPC_FLOAT is a double).
Verified: testsuite x3 + ctest 297 (ASan Debug), RelWithDebInfo full
ctest, lpc-syntax node tests (50).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Reorder sidebar: Driver > CLI > Reference (LPC Language, Apply, EFUN, Concepts)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Fix Docusaurus build: broken links, duplicate routes, and gh-pages CI
- Strip .html from all markdown link targets (51 files) for Docusaurus URL routing
- Add slug: frontmatter to 4 files whose names match their parent directory
(interactive.md, objects.md, README.md, build.md) to prevent Docusaurus's
category-index convention from creating duplicate routes
- Fix one missed .html link in zh-CN/build/index.md
- Move onBrokenMarkdownLinks to markdown.hooks (Docusaurus v4 deprecation)
- Update gh-pages.yml: rename to Docusaurus, use node 22, correct build path
(docs/build instead of docs/.vitepress/dist)
Build now completes with [SUCCESS] and zero warnings or broken links.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* 11-12-adding_nullish_coalescing_operator: 2025-11-12 22:47 - nullish stuff
* 11-12-adding_nullish_coalescing_operator: 2025-11-12 23:28 - adding logical assignment operators
* adding autogen files because grammar has changed
* addressing Codex feedback
* Fix __TREE__ debug output for NODE_NULLISH and NODE_LOGICAL_ASSIGN
The lpc_tree_name array was missing entries for NODE_NULLISH and
NODE_LOGICAL_ASSIGN node types that were added when implementing
the nullish coalescing operator (??). This caused __TREE__ to return
incorrect type names in the debug output.
The fix adds the missing entries "nullish" and "logical assign" to
the lpc_tree_name array at the correct indices to match the parse node
enum definition, allowing the constant_expr.c test to pass.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* adding autogens
---------
Co-authored-by: Claude <noreply@anthropic.com>
* first-class-functions
* first-class-functions
* adding grammar.autogen.cc as requested. did not generate a .h after merging in last update
---------
Co-authored-by: Yucong Sun <1256464+thefallentree@users.noreply.github.com>