Replace the hand-written parser in calc.c with a tree-sitter grammar
(subprojects/rizin-math-parser) and a typed evaluator. The old parser
could only ever produce a ut64 and folded anything it failed to read to
0, which left callers unable to tell a failed expression from one that
evaluated to zero.
Expressions now evaluate to an RzNumValue, a tagged union over ut64,
double, RzBitVector, arbitrary-precision integer and arbitrary-precision
decimal, carrying an RzNumError rather than signalling failure as 0.
Literals keep the width they were written with (5u8, 0xffu128, any width
from 1 to 65536), results that outgrow 64 bits promote to a big number on
their own, and a parse error, division by zero or unresolved identifier
reaches the caller.
rz_num_math() is deprecated. rz_num_math_ut64() keeps its exact behaviour
for callers that want a ut64, and rz_num_math_value() exposes the typed
result. rz_core_math() adds the RzCore-backed form used by the % command,
with rz_core_math_ut64() deprecated alongside it. rz-ax routes through the
typed API, so it prints values at full precision, reports errors on stderr
and exits non-zero. rz_il_lift_num() converts an expression to an
RzILOpPure, so a numeric argument can be lifted instead of pre-evaluated.
Legacy input still works: trailing base suffixes (101b, 35o, 212t), the
trailing-'h' hex form and the k/m/g scale suffixes are all accepted and
warn once, pointing at the 0b/0o/0t prefixes. doc/math.md documents the
language and doc/math-il-lift.md the lift; the grammar, the evaluator,
rz-ax and the % command are covered by unit and db tests.
* Update rz_config list variables to set variables
* Linking error fix
* Update rz_config_get_options in cautocmpl.c
* Update rz_config_get_options in core/tui/config.c
* Test fix
* Assertion error fix
To perform the effect in a delay slot, if the branch was not taken, the
IL, which is already lifted as part of the delay slot instruction, would
explicitly jump to itself again, to execute the effect as normal.
This would create erroneous loop edges in the cfg.
It is actually not necessary to perform this jmp since we already have
the lifted effect and can inline it.
While processing xrefs for marking them as data:
1. classify target using `xref_ref_kind` for data section too, previously it was only classified if target was in exec segment. Which caused false positive when the target was in non-exec section. Happens when the immediate value is small and it points in data section.
2. restrict data block from bleeding into other sections. Currently it correctly caps data block at next "detected" function (or next data) but when the function is not detected yet (like in stripped bins) and the area onward from data ref is empty, the data block bleeds into other sections specifically executable section. This should never happen.
Flags are sorted into the name hashtable with their realnames as well.
Refcounting is used to prevent double-free and similar issues that would
be caused by this.
rz_reg_profile_to_cc() only emitted the first four argument registers
(A0-A3), so architectures that pass more arguments in registers -- the
C6000 EABI uses ten, and x86-64/riscv/ppc all declare more than four --
got a truncated convention. Walk the whole A0-A9 role range, stopping at
the first role the profile leaves undefined, and build the cc string with
RzStrBuf. Covered by a new test_reg unit test.
Co-authored-by agent: Claude/claude-opus-4-8
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
Warning: this also swaps the arguments of the old rz_bv_append() to be
consistend with the new inplace variant.
The reason why the inplace function has the low as the first operand is
that it can be more efficient to append to an existing vector inplace
than to prepend to it. Then, the first argument is being used as the
in-out one in all other inplace functions.
* Don't print meta items which are not at the current seek.
The old code tried (unsuccessfully) to print _any_ meta item _covering_ the seek (ds->at).
There seems to be several bugs getting triggered with that.
One of them giving the behavior of https://github.com/rizinorg/rizin/issues/6556.
If the current seek is in a _data_ region, the disassembler logic doesn't care.
It just assumes that RzAsmOp.size is equivalent to the size of the objects there.
Even though there are only Meta items.
But since some meta items are like 4K bytes, RzAsmOp.size gets
trimmed down.
Anyways, that completely messes up the size calculation (as can be seen in the issue),
and the navigation.
I couldn't figure out where stuff broke.
But the library closes and I have to leave, so I push that.
That "fix" makes it at least behave somewhat consistently.
* Fix leaks
* Fix and add interactive test
- Introduced md1img.h and md1img.c for parsing MediaTek md1img container format.
- Implemented mtk.h and mtk.c for parsing MediaTek GFH firmware images (md1rom).
- Added plugin support for md1img and mtk formats in bin_md1img.c and bin_mtk.c.
- Updated meson.build to include new source files and plugins.
- Enhanced RzBuffer utility with LZMA alone decompression support.
---------
Co-authored-by: Giovanni <561184+wargio@users.noreply.github.com>
* arch/tms320: add TMS320C54x disassembly support
Add a C54x instruction decoder that reuses the shared C55x decode engine
(c55_decode/c55_format) via the C55ArchDesc plug-in interface, rather than
duplicating the matcher/formatter. Disassembly only for now (.lift = NULL).
Engine changes (c55_ir.c/.h):
- add C55ArchDesc.words_le so the decoder can byte-swap the little-endian
16-bit instruction words used by the C54x COFF object format;
- add a self-contained C54x memory-operand renderer (direct @dma, MMR,
indirect *ARx with all post-modify modes, *ARx(lk) const-index, *(lk)
ABS16 absolute and circular '%' addressing) and bare-hex immediates;
- add C55Operand.circular for the '%' suffix and C55Operand.space_join
for the space-separated second half of a C54x parallel instruction;
- extend the data-memory operand-field analysis (register, base pointer,
displacement, direction, referenced size) to the LOAD/STORE op types the
C54x ld/st family uses, in addition to the C55x MOV form.
The C54x decoder (isa/tms320/c54x/c54x.c) covers the complete documented
instruction set - all 117 mnemonics of the SPRU172 opcode map, in every
documented encoding form:
- load/store/move, integer and logical ALU ops in every addressing form
(Smem, #lk, dual-accumulator, Xmem/Ymem, TS/ASM/SHIFT-shifted, the
shift-by-16 and #lk,16 long-immediate forms, and the two-word
Smem,SHIFT form whose operation selector lives in the second word);
- the full multiply/MAC family: Smem, #lk, program-memory, squaring,
multiply-by-A, signed-unsigned and the dual-operand MAC[R]/MAS[R]
Xmem,Ymem forms;
- the parallel (dual-operation) class rendered "op1 .. || op2 .." -
ST||ADD/SUB/LD/MPY/MAC[R]/MAS[R], ST||LD T and LD||MAC[R]/MAS[R];
- double/long-word (Lmem) add/subtract, the unary accumulator ops
(exp/norm/abs/neg/rnd/sat/min/max/rol/ror/sftc/cmpl/...);
- control flow with the separate delayed (bd/calld/bcd/banzd/fcalad/...)
variants, conditional return/execute (rc[d]/xc) and the multi-condition
"tc, c"-style combinable condition fields, repeats (incl. rpt #lk),
conditional stores, I/O port access, status-bit set/clear and the
non-linear idle encoding.
Operands resolve to their architectural names - the full memory-mapped
register file (AR0-AR7, the accumulator AL/AH/AG/BL/BH/BG halves, T, TRN,
SP, BK, BRC/RSA/REA, IMR/IFR, PMST, XPC), the ST0/ST1 status bits and the
named condition codes; the memory-mapped-register operand is kept single
word (its long-offset modes are not legal). The analyzer classifies every
instruction (op->type, op->id), resolves branch/call targets and the stack
effect of calls/returns/pushes, and exposes operand details: the register,
base pointer, displacement and access direction of data-memory loads and
stores, and the target register of indirect branches/calls.
All encodings were verified byte-exact against the TI asm500 assembler,
and every decoded instruction re-assembles to an identical encoding (a
full-opcode-space disassemble/reassemble round-trip is stable). A 297-case
disasm test suite and an analysis test suite (opcode classification, branch
and call targets, stack effects, memory-operand fields, data-immediate values, the register
profile, named instruction ids and COFF binary-fixture function discovery)
are added, and the real-world emulateme C54x .text decodes cleanly.
* arch/tms320: add TMS320C54x RzIL lifting
Lift the C54x integer core to RzIL so emulation and IL-based analysis work
for C54x as they already do for C55x/C55x+.
- Register profile: C54x previously fell through to the C64x profile
(a0-a31, =PC pce1), wrong for the A/B accumulator core. Add a proper
C54x profile: the two 40-bit accumulators A/B (with the L/H 16-bit and
G 8-bit guard slices overlapping their parent), AR0-AR7, T/TRN, SP, DP,
BK, ST0/ST1/PMST, BRC/RSA/REA, IMR/IFR, XPC and a 24-bit PC.
- IL VM config: tms320_c54x_il_config() binds the canonical registers; the
accumulator slices stay unbound, the lifter expresses them as bit-slices
of A/B so they never desynchronise.
- Lifter (C55ArchDesc::lift hook, dispatched by c55_lift): the no-shift
forms of LD/LDU/LDR/LDM, ADD/SUB/AND/OR/XOR, STL/STH/STLM/STM, the mvd*
memory-to-memory moves, the DLD/DST 32-bit double-word load/store (high
word at the lower address), PSHM/POPM and RET. Shift/round/saturate
variants are left unlifted (their shift count is carried only as a
display string); the engine's generic EA/read/write/post-modify helpers
are reused for the addressing modes.
Tested via two new RzIL VM blocks in test/db/rzil/tms320: a register/
immediate/memory execute test, and an end-to-end emulation of the
emulateme binary's _decrypt (a UART hex-writer) showing the IL VM emits
the hex digits and advances the write position.
---------
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* Cast operands for AND and OR instructions to the correct width
* Add missing operand casts for SBB and ADC.
* Add flawed instructions to asm tests
---------
Co-authored-by: Dhruv Maroo <dhruvmaru007@gmail.com>
* Add reliable http:// test
* REUSE.toml: Add `test/www/**` entry
* Use `cwd` instead to work around old http.server in Python 3.6
* Move test to `not-windows-any`
* NetBSD: Add `python3` symbolic link
* Prevent test from running on woodpecker