Expose the typeclasses of types through the new `tk` commands:
- `tk <type>` shows the typeclass of a type;
- `tkl` lists the available typeclasses, while `tkll` (verbose), `tklt`
(table) and `tklj` (JSON) additionally list the types belonging to each
typeclass;
- `tks <type> <typeclass>` sets the typeclass of a type.
A new rz_base_type_set_typeclass() API backs the `tks` command. Since a
typedef without its own typeclass inherits the one of the type it points
to, setting the typeclass of an atomic type is automatically reflected on
the typedefs resolving to it.
The `tk` help also explains what typeclasses are and lists the available
ones.
Closes#3371
An enum type hint set on an operand (ahie) only replaced the immediate when its
value matched an enum member exactly, via rz_type_db_enum_member_by_val(). A
value that is the bitwise OR of several flag members -- e.g. the access(2) mode
R_OK|W_OK == 6 from the issue -- matched no single member and was left as a raw
number.
replace_enum_hint() in rz_parse now falls back to rz_type_db_enum_get_bitfield()
when there is no exact member, so the value is rendered as the OR of the
matching members, e.g. "access_def.W_OK | access_def.R_OK".
rz_type_db_enum_get_bitfield() is reworked along the way: it was unused and
buggy (capped at 32 bits, reused a stale match for bits without a member, and
emitted a "0x.. : " debug prefix). It now walks all 64 bits, returns the
matching members qualified with the enum name and joined by " | ", and returns
NULL when the value is 0, the type is not an enum, or any set bit has no member
(so a value that is not cleanly a combination of flags stays a number).
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* Fix revert if adding a node failed.
* Reduce allocations by keeping only a single RzGraphEdge object per edge around.
* Add benchmark for graph deletion and addition of nodes/edges
* Use rz_pvector_remove_at_unsorted to save some runtime.
* Use realloc and memmove for matrix graphs on capacity increase.
* Add helper to determine memory usage.
* Add benchmark
* Revert matrix capacity extension to simple and jsut as fast loop.
* Missing type annotations
Introduce rz_type_db_rename_base_type() and expose it through the new
`tr` command to rename a base type. In addition to renaming the type
itself, every reference to it is updated so the analysis state stays
consistent after the rename:
- other base types and function types (callables) that use the type,
including self-references such as a linked-list struct that contains a
pointer to itself;
- the `pf` formats that mention the type, both the format stored under
the type's own name and the "(name)" references inside other stored
formats;
- the types of analysis global variables (as reported by `avg`), the
function signatures (return type and arguments) and the function local
variables.
The whole update is orchestrated by rz_core_types_rename(), and the
recursive RzType reference renamer is exposed as
rz_type_rename_references() so it can be reused for type usages that live
outside the type database.
Closes#1078
The C grammar already exposes a struct/union member's bitfield width as a
"bitfield_clause" node and the parser already recognized "int a : 4;", but
it threw the width away (the member was stored with size 0 and a FIXME). As
a result "tc" rendered the member as a plain "int a;", "ts" produced a
full-int "pf" format ("pf d4d4d4 a b c") and "tp" mis-read every field as a
whole integer.
RzTypeStructMember and RzTypeUnionMember already carry a "size" field
documented as "in bits"; it is now used as the bitfield width, where 0 means
the member is not a bitfield:
- The struct and union member parsers store the parsed bit width on the
member instead of discarding it.
- The pretty printer emits " : <width>" before the member's trailing ';',
so "tc"/"tcd" round-trip "struct qwe { int a : 4; int b : 16; int c : 3; }".
- rz_base_type_as_format() (used by "ts"/"tp") now emits the pf packed-bits
spec ":N" for a bitfield member, with the bit order matching the target
endianness ("<" little-endian / ">" big-endian), instead of the member
type's full format. "ts qwe" becomes 'pf ":4<:16<:3< a b c"' and "tp"
unpacks the fields correctly (a=5, b=0x1234, c=0 for 0x00012345 on x86).
Serialization keeps the width in the 3rd comma-field of the existing
"struct.<name>.<member>" / "union.<name>.<member>" value, i.e.
"type,offset,bitsize". That field was already written (always as 0) and
ignored on load, so the on-disk layout is unchanged and only its meaning is
now honored. The project version is bumped to 24 with an additive no-op
migration: an absent or 0 bitsize deserializes as a regular, non-bitfield
member.
Tests: a unit test parses a bitfield struct and union and checks the member
widths, the rendered string and the "pf" format; a db/cmd/types test covers
the full "td"/"tcd"/"ts"/"tp" flow; a project-migration test with a new
v23-bitfield.rzdb fixture covers v23->v24, and the "migrate info" db test is
updated for the new version.
The PDB type parser (librz/arch/pdb_process.c) still drops bitfield members
because their LF_BITFIELD member type resolves to NULL; wiring LF_BITFIELD up
to the new member "size" field, with tests in test/unit/test_pdb.c, is left
as a follow-up.
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* Update Capstone and add m68k ColdFire support
* Fix test 'core regs linux m68k'
* Add M68K ColdFire integration coverage
* Update Capstone v6 to version 6.0.0-Alpha9
* Refactor CPU mode detection to use case-insensitive string comparison
* Update Capstone next revision to 3df6ff0
The ELF section classifier in sections_obj() only marked a section as
data when its name contained "data". As a result .dynstr (type
SHT_STRTAB, SHF_ALLOC) was flagged neither as data nor as containing
strings, so is_data_section() rejected it. Two problems followed:
- the string search (iz, and the default AUTO scan used by -A) skipped
.dynstr, so its strings were never listed (only izz, which scans the
whole file regardless of section flags, showed them);
- because no RZ_META_TYPE_STRING metadata was applied over the region,
the disassembler rendered the NUL-terminated symbol names as code,
e.g. on MIPS "_GLOBAL_OFFSET_TABLE_" decoded to bgtzl/ldr/jalx/...
Mark a section as containing strings when it is mapped into memory
(SHF_ALLOC) and is either a string table (SHT_STRTAB) or carries the
explicit SHF_STRINGS flag. The SHF_ALLOC restriction keeps the loaded
string tables (.dynstr) while leaving the non-allocated .strtab and
.shstrtab to izz, matching the "iz lists the loaded image" semantics.
This is architecture independent; MIPS was simply where the bad
disassembly was first noticed.
Closes#5182
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
C23 lets an enum fix its underlying type, e.g. "enum E : long long { ... }".
The C grammar already exposes it as the "underlying_type" field of an
enum_specifier, and RzBaseType already has a "type" slot documented as
used by enums, but the parser ignored the field and always left it NULL.
- parse_enum_node() now reads the "underlying_type" field and stores the
parsed type on RzBaseType::type (reusing parse_type_node_single(), so
primitive, sized and typedef'd integer types are all handled). Classic
enums keep a NULL underlying type.
- The pretty printer emits " : <type>" between the enum name and its body
when an underlying type is present, so "tc"/"tcd"/"tec" round-trip it.
- enum_bitsize() now derives the width from the underlying type instead of
the hardcoded 32-bit default (resolving the long-standing FIXME); it
still falls back to 32 for a classic enum.
Single-token underlying types (int, uint64_t, char, ...) work end-to-end
with the bundled grammar revision: the existing "enhanced enum" db test is
updated to round-trip "enum v : int" and a unit test covers
"enum EU : uint64_t". Multi-word underlying types (long long, unsigned int)
are added as BROKEN db tests; the bundled grammar parses them to an ERROR
node and drops the underlying type, so these tests fail for now and will
pass once rizin-grammar-c accepts sized type specifiers as the enum
underlying type. No further rizin change is needed for that step.
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
Close two long-standing issues against the static horizontal histogram
(`p==`, `p==e`, `p==0` and the rest of the `p==X` family). Add the
`scr.hist.width` / `scr.hist.height` config options (closes#6372),
both clamped via RZ_MIN() against the current terminal size with a
defensive sanity floor. Make the Y-axis ruler context-aware (closes
#5290) via new `value_min` / `value_max` / `value_unit` fields on
`RzHistogramOptions` and rescale each datum from its ut8 storage into
the vmin..vmax display range before thresholding, which also fixes
cer-0's follow-up about the chart's top rows staying blank for
low-range data.
Tighten the look in the same step. Y-axis labels are now sparse (top,
~25 %, ~50 %, ~75 %, bottom) instead of one per row, in the style of
tokio-console / Granite, with an inclusive label range (top = vmax,
bottom = vmin) decoupled from the threshold formula so the existing
`_` baseline behaviour still appears. The bottom of the chart grows
an X-axis ruler with `^` tick markers and `0x...` start / middle /
end offsets (driven by a new `opts->blocksize` field). Two more
options, `value_scale` and `value_precision`, let callers format
ruler labels as fractional values; the `p==e` (entropy) command opts
in with `value_max = 1, value_precision = 2` so the ruler reads
0.00..1.00 by default, matching Shannon-entropy convention.
A new `opts->data_f` field lets callers feed double-precision data
directly into the renderer (and engages double-precision arithmetic
for the row thresholds), so entropy keeps full precision end-to-end
instead of losing ~1/255 of resolution to the `(ut8)(255 * fraction)`
quantisation step. This captures the precision win sketched in PR
#6355 without introducing a separate `rz_histogram_horizontal_f64()`
twin function. Ten unit tests in `test/unit/test_cons_histogram.c`
(seven new) cover sparse labels, fractional labels, the X-axis offset
ruler, the cer-0 regression, `opts->cols`, the legacy 0..255 ruler,
and the fp data path; all pass.
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
With bin.demangle=true the call/jump operand is replaced by the full
demangled symbol name and then colorized token-by-token before being
filtered. For long C++ symbols (e.g. libc++ STL names) the colorized
operand easily exceeds the 1024-byte operand buffer (ds->str): in
truecolor each token is wrapped in a ~19-byte "\x1b[38;2;R;G;Bm" escape,
so a ~110-char name grows past 1700 bytes.
The generic filter() copied that operand with
strncpy(str, data, len);
which has two problems when strlen(data) >= len:
- it does not NUL-terminate, so reading str ran past the buffer into
the adjacent strsub[] buffer, printing stale bytes (the "call 0x..."
bleed); and
- the byte-bounded cut can land in the middle of a color escape,
leaving a partial "\x1b[38;2;.." that, followed by the bled bytes,
forms a complete but bogus "\x1b[..c" sequence. Terminals read that
as a Device-Attributes query and reply with e.g. "62;4c", which is
the garbage the reporter saw at the prompt.
This only triggers with color enabled and a symbol long enough that its
colorized form overflows the buffer, which is why it showed up on some
x86 files only and had no portable reproducer (no-color = 149 bytes,
16-color = 763, 256-color = 1170, truecolor = 1713 for the sample
symbol; the threshold is 1024).
Add test/unit/test_parse.c driving rz_parse_filter() with over-long
colored operands (a deterministic case whose cut lands inside an escape,
and a realistic libc++-style colorized operand) plus a short-operand
passthrough case. The two overflow tests fail on the old strncpy and
pass with the fix.
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
format.c can already turn an RzType into a pf format string. Add the
inverse in the same file, next to its counterparts:
rz_type_format_to_c_declaration() parses a pf format string with
rz_pf_parse() and emits an equivalent C struct/union declaration built
from the standard fixed-width types (uint8_t, int32_t, float, ...).
The conversion is structural: it consumes only the parsed RzPfFormat
(field kinds, widths, array counts, pointer flags), never a byte buffer,
so it runs without any target data. Every field kind the engine produces
is mapped; a handful of specifiers with no exact static C form are mapped
best-effort (documented inline): @N alignment is dropped, an
unknown-length inline z string becomes char *, LEB128 widens to its
largest decoded integer, and ?(Name) / E(Name) emit struct/enum
references that must themselves be defined to parse.
Added new 'tdf' command that exposes that conversion to the user
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
Most of the Windows types reported in issue #3728 were stub entries of
the form
NAME=type
type.NAME=p
i.e. opaque atomic pointers with no actual body. rz-ghidra rejects
these as 'Invalid atomic type'. Convert them to proper typedef entries
backed by either real struct definitions (for well-known types) or
forward-declared opaque structs (for architecture-specific ones like
CONTEXT and KNONVOLATILE_CONTEXT_POINTERS, whose layout varies across
x86/x64/ARM/ARM64 and which the shared types-windows.sdb.txt cannot
commit to).
Specifically, this commit:
* converts the stubs for PSID, LPOLESTR, LPCLSID, PCONTEXT,
PKNONVOLATILE_CONTEXT_POINTERS, LPSYSTEM_INFO, LPWIN32_FIND_DATAW,
PCACTCTXW, PSID_AND_ATTRIBUTES, PLUID_AND_ATTRIBUTES, PTRACEHANDLE
and the duplicate LPMSG/PLUID into proper typedefs;
* adds new top-level typedefs for CLSID, OLECHAR, LPCOLESTR
(corrected to const OLECHAR *), REFCLSID, LPBC, MSG, PMSG, NPMSG,
LPMSG, SYSTEM_INFO, WIN32_FIND_DATAW, PWIN32_FIND_DATAW,
OSVERSIONINFOW, POSVERSIONINFOW, LPOSVERSIONINFOW,
RTL_OSVERSIONINFOW, PRTL_OSVERSIONINFOW, OSVERSIONINFOEXW (plus its
POSVERSIONINFOEXW/LPOSVERSIONINFOEXW siblings), ACTCTXW, PACTCTXW,
LUID, PLUID, LUID_AND_ATTRIBUTES and SID_AND_ATTRIBUTES;
* adds struct bodies for tagMSG, _OSVERSIONINFOW, _OSVERSIONINFOEXW,
_SYSTEM_INFO (with the union flattened to wProcessorArchitecture +
wReserved per modern usage), _WIN32_FIND_DATAW, tagACTCTXW,
_SID_AND_ATTRIBUTES, _LUID, _LUID_AND_ATTRIBUTES and the empty
opaque tags _CONTEXT, _KNONVOLATILE_CONTEXT_POINTERS and _IBindCtx;
* fixes a small typo in struct._EVENT_DATA_DESCRIPTOR where the
field list said 'Usize' while the per-field entry was 'Size'.
Member layouts were taken from the Microsoft Win32 documentation
(winnt.h, sysinfoapi.h, winuser.h, evntrace.h, minwinbase.h,
winbase.h) and cross-checked with the Wine include/winnt.h,
wtypes.idl and guiddef.h headers.
Offsets follow the 32-bit Windows ABI used by the rest of this file
(this matches the existing _OBJECT_ATTRIBUTES, _FILETIME etc. layouts).
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
When type_to_format / type_to_format_pair encounter a pointer field whose
pointee is reachable only through a typedef chain that ends at the
atomic 'void' or 'char' (e.g. PVOID -> VOID -> void, LPSTR -> CHAR ->
char, HANDLE -> ... -> void in some platform headers), they emit a bare
'*' and recurse into the pointee. The recursion produces nothing for
'void' (no format) and emits the legacy 'c' for 'char', so the
generated pf string ends up with either an orphan trailing/internal '*'
or the sequence '*c', which the new pf parser introduced in the
recent rewrite rejects: it requires '*' to be followed by a complete
dereferenceable spec ('z', a sized integer like 'd4'/'x2'/'u8', or
'?'). The result is parser warnings and dropped fields during 'tp':
rizin -k windows -c 'tp _OBJECT_ATTRIBUTES'
WARNING: pf: unknown specifier '*' at position 4, skipping
WARNING: pf: unknown specifier '*' at position 4, skipping
[only 4 of 6 fields rendered]
rizin -k windows -c 'tp _SYSTEM_INFO'
[11 fields collapse to 4]
rizin -k windows -c 'tp _STARTUPINFOA'
[three LPSTR fields render as '*c' which the parser cannot
usefully follow]
The top-level rz_type_as_format() already special-cases 'void *',
'char *', and callable pointers, mapping them to 'p' and 'z'. But the
inner walkers do not, because rz_type_is_void_ptr / rz_type_is_char_ptr
compare the literal identifier name and so do not see through typedefs
like VOID->void or CHAR->char.
Fix it by adding ptr_pointee_resolves_to(), a small static helper in
format.c that walks the typedb to find the canonical atomic name for
the pointee (bounded depth so a circular typedef cannot send the
resolver into an infinite loop), and using it in both POINTER branches
before the '*'+recurse fallback. Pointers whose pointee resolves to
'void' (or a void-aliased typedef) now emit 'p'; pointers whose pointee
resolves to 'char' (or a char-aliased typedef) emit 'z'. Everything
else continues to emit '*<inner>' unchanged, so single-level
LPBYTE -> BYTE -> unsigned char still renders as the perfectly valid
'*x1', and pointer-to-struct '*?' chains are untouched.
After the fix, the same upstream structs produce well-formed pf
strings:
ts _OBJECT_ATTRIBUTES -> 'x8p**x1x8pp ...' (two PVOID -> pp)
ts _SYSTEM_INFO -> 'x2x2d4ppx8d4d4d4x2x2 ...' (two LPVOID -> pp)
ts _STARTUPINFOA -> 'd4zzzd4...*x1ppp ...' (three LPSTR -> zzz;
LPBYTE still '*x1')
Add a regression test in test/db/cmd/cmd_pf that exercises both shapes
via 'ts _SECURITY_ATTRIBUTES' (single LPVOID) and 'ts _STARTUPINFOA'
(three LPSTR + one LPBYTE). Without this commit the test catches the
bug -- the LPVOID field becomes orphan '*' inside the format and the
LPSTR fields render as '*c'; with the commit both render cleanly as
documented above.
Also update one existing EXPECT in test/db/cmd/types: the 'td with
comments' test had encoded the legacy 'pf "d4[5]*c b foo"' shape for
'char *foo[5]'. With the fix this becomes 'pf "d4[5]z b foo"', which
is both well-formed under the new DSL and a more accurate description
(an array of zstrings rather than an array of pointers to a single
char).
Co-authored-by: Anton Kochkov <anton.kockov@gmail.com>
* librz/type: rewrite pf format parser
Replace the legacy print-format engine in librz/type/format.c with a
clean three-stage pipeline (parse -> read -> render) and a properly
typed DSL.
The new pipeline:
rz_pf_parse(str) source string -> RzPfFormat
rz_pf_read(fmt, buf, len, ctx) RzPfFormat -> RzPfValue[]
rz_pf_render(vals, mode, opts) RzPfValue[] -> printable string
rz_pf_format() one-shot for all of the above
Public API lives in <rz_pf.h> (pulled in transitively via <rz_type.h>).
DSL improvements over the legacy scheme:
* Sized integers with explicit endianness via case:
x4 d4 u4 o4 b4 (LE) vs X4 D4 U4 O4 B4 (BE).
* Sized floats: f2 / f4 / f8 (and BE F2 / F4 / F8).
* Encoding-aware strings: z(utf8) / z(utf16le) / z(utf16be) /
z(utf32le) / z(utf32be) / z(latin1) / z(ebcdic).
* Length-prefixed strings: z[N] reads a N-byte LE prefix then body.
* Parameterized timestamps: t(unix32) / t(unix64) / t(unixms) /
t(unixus) / t(unixns) / t(filetime) / t(dos) / t(hfs) /
t(oletime) / t(webkit) / t(cocoa). ntfs alias maps to filetime.
* GUID: G(ms) (Microsoft mixed-endian, default), G(be), G(le).
* Inline bitfields: B4(FLAG_A=1,FLAG_B=2,FLAG_C=4).
* Alignment: @N advances to next multiple of N.
* Bit fields: :N with optional < (LSB-first) or > (MSB-first).
* Length-by-reference arrays: [@earlier_field]T.
* TLV records: V(t=u1,l=u2,e=le,h=v|l|a,d=table). Dispatch table
uses tlv.<name>.<hex-tag> in the typedb formats hash.
Architecture highlights:
* Per-instance ReadState (bit_cursor + sibling lookup window).
The bit cursor snap-flushes to the next byte when a non-bit
field follows mid-byte.
* Recursion bound via RzPfCtx::max_depth (default 32).
* Five render modes: text (default), json, cstruct, quiet, dot.
The dot renderer emits a column-aligned record with offset/type/
name/value rows, suitable for direct rendering through Graphviz.
* Positioned diagnostics: RzPfError carries severity, category,
column position, and a human-readable message. Each is also
emitted via RZ_LOG_WARN for backward compatibility.
* Optional RzPfPalette and RzPfRenderOpts threaded through render
for inline ANSI colorisation. NULL palette = canonical text.
* Pointer dereference via RzPfCtx::read_at callback.
The integration test expectations in test/db/cmd/cmd_pf* are
updated in the same commit to match the new text-output format and
the new DOT renderer's column-aligned record layout. Bundling these
keeps the test suite green at this commit.
A transitional rz_type_format_data() shim at the bottom of
pf_parser.c forwards legacy callers (cprint.c, disasm.c) to the new
API; it is removed in the librz/core migration commit.
* librz/arch/types: migrate type DB to new pf format codes
Map every legacy single-character format code in the bundled type
SDB files to its new-DSL equivalent:
b / C -> x1 (1-byte hex)
c -> c (signed char, unchanged)
w -> x2 (2-byte hex LE)
i -> d4 (signed 32-bit decimal LE)
d -> x4 (4-byte hex LE)
q -> x8 (8-byte hex LE)
f -> f4 (32-bit float LE)
F -> F8 (64-bit float BE)
Z -> z(utf16le)
x -> x4
o -> o4
X -> r (hexdump)
n / N -> d4 / u4
Also add new typedef-only types for explicit display formats:
bin{8,16,32,64}_t -> b{1,2,4,8} (binary)
hex{8,16,32,64}_t -> x{1,2,4,8} (hex)
oct{8,16,32,64}_t -> o{1,2,4,8} (octal)
These let users opt into a display representation per field without
changing the underlying scalar size.
Update the corresponding integration test expectations in
test/db/cmd/cmd_avg and test/db/cmd/types: the 'tu', 'tuc', 'tp',
and 'avgp' commands report formats via the SDB, so the migrated
codes appear in their text output (e.g. 'pf 0f4x4 a b' instead of
the legacy 'pf 0fd a b' for a union of float + int).
* librz/core: rewrite pf integration on the new parser API
Remove the legacy rz_type_format_data() bridge and rewrite every
caller to use the new <rz_pf.h> API directly.
librz/core/cprint.c
core_print_format() constructs an RzPfCtx populated from the
RzCore (typedb, big_endian, bits, max_depth) and calls
rz_pf_format() in one shot. The bridge between RzCore I/O and the
pf reader is cprint_pf_read_at(), which forwards to rz_io_nread_at.
The bitmask-to-RzPfMode mapping is cprint_pf_mode().
For DOT mode the format name is passed via RzPfRenderOpts::graph_label
so the dot renderer uses it as the top-level record label.
When scr.color > 0 in TEXT mode, an RzPfPalette is populated from
the active color theme via RzConsPrintablePalette so `pf` output
blends with the rest of Rizin's UI and respects user theme
choices (eco / ec). The palette slots map to theme fields as:
pf field offset <- pal->offset (address column)
pf field name <- pal->fname (symbolic name)
pf endian marker <- pal->meta (metadata tag)
pf hex/number literal <- pal->num (numeric literal)
pf typedb label <- pal->flag (resolved symbol)
reset <- pal->reset
Each slot falls back to its previous hardcoded ANSI escape if the
theme field is unset, preserving prior behaviour for stripped-down
contexts. Other modes (json/cstruct/quiet/dot) ignore the palette.
Closes rizinorg/rizin#782.
librz/core/disasm.c
RZ_META_TYPE_FORMAT case rewritten to resolve via rz_pf_resolve_name
and decode via rz_pf_format() directly. Palette is supplied from
the same theme palette when ds->show_color is true so that
pd @ <struct> blends with the surrounding listing.
librz/include/rz_type.h
Drop the rz_type_format_data() forward declaration. Callers must
use the public <rz_pf.h> surface instead.
librz/core/cmd_descs/cmd_print.yaml
Rewrite the pf command help for the new DSL: sized integers
(case-discriminated endian), special scalars (G GUID, V TLV,
: bits, @ align), strings with encodings, parameterized timestamps,
typed composites (E enum, B bitfield, ? struct), DSL extensions
(length-prefix strings, length-by-reference arrays, inline
bitfields), skip/repeat/pointers, example invocations.
cmd_descs.c is regenerated from the YAML.
* test: pf parser unit and DSL-extension integration tests
Add fresh test coverage for the new pf parser. These are purely
additive: existing integration tests were updated in the parser
and SDB-migration commits so the suite stayed green throughout.
test/unit/test_pf.c (new, 131 tests)
Field-size tables (every fixed-size type) and ctype mapping, every
parse path (sized integers / floats / strings / timestamps /
pointers / structs / enums / bitfields / GUIDs / TLV / alignment /
bits / arrays), every read path (with both fixed and length-by-
reference array counts), every render mode (text / json / cstruct /
quiet / dot), every DSL extension (align, bits MSB/LSB, GUID layout,
length-ref array, length-prefix string, inline bitfield, TLV with
dispatch table), pointer dereference via an in-memory I/O callback,
recursion safety on self-referential structs, lifecycle / NULL
safety, diagnostics (positioned errors, all categories, caret-line
formatter, source capture, verbose parse), ambiguity (each
superficially similar DSL form parses unambiguously), the palette /
colorisation API (issue #782) including a regression test surfaced
by the property-based test harness for input bytes containing ESC,
and DOT-mode coverage (column-aligned layout, single field,
empty/filtered, timestamp values, TLV records).
test/db/cmd/cmd_pf_dsl (new)
8 integration tests targeting the new DSL extensions specifically:
@N alignment, :N bits in MSB and LSB ordering, default-layout GUID,
z[N] length-prefixed string, B4(K=V) inline bitfield, bare V TLV
with defaults, V(t=u2,l=u2,e=be) configured TLV.
test/db/cmd/types_format (new)
Verifies that the new bin{N}_t / hex{N}_t / oct{N}_t typedefs apply
the expected display representation when used inside struct fields
via the 'tp' command, demonstrating the per-base rendering
(binary / decimal / octal / hex).
* doc: pf DSL reference and librz/type README
doc/pf.md (new)
User-facing reference for the pf format DSL covering quick
examples, spec grammar (sized integers with case-discriminated
endian, floats, strings with encodings and length prefixes,
parameterized timestamps, composites -- struct ?, enum E, bitfield B
inline and typed, GUID G, TLV V, raw hexdump r), repetition and
arrays ([N], [@name], {N}, leading 0 for union), padding and
alignment (. skip, @N align, :N bits MSB / LSB), pointer
dereference (*<type>), names grammar (plain vs (typename)name),
output modes (text, json, cstruct, quiet, dot), and diagnostics
(severity / category / position / caret-line formatter; backward-
compatible RZ_LOG_WARN channel).
librz/type/README.md (new)
Subsystem architecture doc for librz/type. Covers public headers,
file map, the pf three-stage pipeline (parse -> read -> render) with
a diagram and entry-point table, per-stage explanations (parser
shape, ReadState scoping rules including bit cursor and sibling
lookup window and snap-flush rule, render mode dispatch including
the DOT column-aligned layout and the graph_label parameter),
typedb integration and recursion bound, error-reporting model,
test coverage summary, and pointer to doc/pf.md as the user-facing
reference.
* librz/type: remove legacy DSL conversion shim
All in-tree callers -- librz/bin/d/, librz/bin/format/, and the
integration tests in test/db/cmd/ -- have been migrated to the new
DSL in the earlier commits of this series. The compatibility shim
that translated bare legacy specifiers (x -> x4, b -> x1, w -> x2,
nN -> uN, NN -> UN, etc.) into the new DSL at the entry of
rz_pf_parse() is no longer load-bearing; this commit deletes it.
- librz/type/pf_parser.c: drop the 267-line
convert_legacy_to_new_dsl() function, the WARN_LEGACY macro,
and the per-parse allocation + free of the converted string.
rz_pf_parse() now walks the caller's spec directly.
- test/db/cmd/cmd_pf,cmd_pf2,cmd_pf_write,cmd_pfd,cmd_pf_new,
metadata: migrate legacy specifiers in CMDS blocks to the new
DSL (b -> x1, w -> x2, q -> x8, i -> d4, bare x -> x4, etc.)
and regenerate EXPECT blocks against the new render where
affected.
The parser now only accepts the new DSL. Any caller still passing
bare legacy specifiers will get "unknown specifier" warnings.
* librz/core,type: add scr.pf.short for delta-offset rendering
When scr.pf.short is enabled, MODE_TEXT rendering shows offsets as
deltas (+<n>) from the format's base address instead of the absolute
hex address. Useful for self-contained struct dumps where the
absolute load address is noise:
$ rizin -e scr.pf.short=true -qc 'wx 00007a452a4b9a02
pf fcb1d4 a b c d' =
0 : a = 4000 [LE]
+4 : b = '*'
+5 : c = 0b0100_1011 [LE]
+6 : d = 666 [LE]
Nested structs reset their delta from the same base, so a child of
`head` at offset +8 still prints as +0 in its own struct block.
The base is derived from the offset of the first top-level value;
callers can override it via RzPfRenderOpts::base_offset.
Includes 5 integration tests in test/db/cmd/cmd_pf_short covering
basic struct, off-mode parity (control), nested struct, concatenated
TLV records, and @-offset usage.
* librz/type,test,doc: add v(N) bitvector type for forensics use
Adds a new specifier `v(N)` to the pf DSL that reads N individual
bits (1..4096) from `ceil(N/8)` bytes and exposes them as N
separate 0/1 scalars. Forensics targets: page-frame bitmaps, NTFS
$Bitmap clusters, ext4 block/inode bitmaps, ELF DT_FLAGS_1, PE
characteristics, ACL bitmasks -- anywhere you want to *see* a
bitmap rather than collapse it to a hex number.
Grammar:
v(N) N-bit bitvector, default MSB-first per byte
v(N,lsb) LSB-first per byte (Intel order)
v(N,msb) explicit MSB-first (DWARF / network order)
Rendering:
text: [ 1 0 1 0 1 0 1 1 | 1 1 0 0 ] (12-bit)
json: {"bit_width":12,"value":"101010111100"}
quiet: 1 0 1 0 1 0 1 1 1 1 0 0
Bitvec fields do not participate in the packed-bit cursor used by
:N; each v(N) reads whole bytes and stands alone. A v(N) field
adjacent to a :N field flushes the partial bit cursor first.
Width is clamped to [1, 4096] with a RANGE diagnostic. The reader
consumes exactly ceil(N/8) bytes, leaving the cursor positioned for
the next field. Verified by the existing tail-field test pattern.
Tests:
test/unit/test_pf.c +10 unit tests (146 total OK)
test/db/cmd/cmd_pf_new +9 cmd tests (440 OK / 9 BR / 0 XX / 4 FX
across all 19 pf suites)
property harness +7 properties (140 unique props, 1000
trials each, 0 failures)
Docs:
doc/pf.md bitvector section with examples
librz/type/README.md cursor-interaction note for v(N) vs :N
* librz/core,type,test: address PR CI feedback for pf rewrite
Five fixes prompted by CI feedback and reviewer comments on the pf
rewrite series:
* librz/core/cmd_descs/cmd_print.yaml: two help-text lines exceeded the
yamllint 120-char ceiling (n1/n2/n4/n8 comment and the deprecation
notes entry). Both now use the same folded-scalar (comment: >) form
already used in cmd_analysis.yaml and cmd_descs.yaml. cmd_descs.c is
regenerated.
* test/db/cmd/cmd_pf_write: drop the BROKEN 'pf xxd print' test. It
relied on '.pf*' (run pf output as rizin commands) which the new
parser no longer emits -- analogous to the '.pfw' removal done
earlier. The companion 'pf xxd print happy' test exercises the same
path through 'pf.' and stays.
* librz/type/pf_parser.c: thread the RzTypeDB through render_cstruct
and render_dot (and their inner helpers) so RZ_PF_ENUM resolution
works in pfc and pfd mode the same way it already does in text mode.
Without this, 'pfc elf_header' emitted '/* 0x00000002 */' and 'pfd'
emitted '|value|0x0000ffff|'; with it the comments and cells carry
the symbolic name (ELFCLASS64, ET_HIPROC, etc).
scalar_text() RZ_PF_BITFIELD: when no inline B4(K=V,...) flags are
present but val->type_name names a typedb enum, walk that enum and
treat each case as a settable bit name. This makes the typedb form
'B (pe_characteristics) flags' decode to
'0x00008140 : IMAGE_DLLCHARACTERISTICS_DYNAMIC_BASE | ...' instead
of just the raw hex -- previously the bitname decoding only worked
on inline B4(...) forms.
* librz/core/cprint.c: cprint_pf_mode() had no mapping for
RZ_PRINT_VALUE, so 'pfv' fell through to TEXT mode and emitted
'0x00000004 = 0x11111111 [LE]' instead of the bare value '0x11111111'
expected by db/esil/arm_16 and existing pfv users. Added an explicit
RZ_PRINT_VALUE -> RZ_PF_MODE_QUIET mapping; QUIET already emits one
bare scalar per field, which is exactly the pfv contract.
* test EXPECTs: cmd_pfd unbreaks its bitfield case (also fixes the
data buffer -- the original used 'wx 0x00008140' which is parsed
byte-by-byte and yields LE u32 0x40810000, not 0x00008140, so no
flags matched). cmd_pf 'PE test' picks up the now-decoded
characteristics and dllCharacteristics bitfields. cmd_pf 'Print
value only' updates to the bare-value pfv contract. test/db/cmd/print
'elf 64bit ls pfc.elf_header' picks up the decoded enum symbols.
Also picks up clang-format-20 reflows in cprint.c, cmd_print.c,
pf_parser.c, and rz_pf.h that the CI's clang-format check flagged.
Also bumps the offset-delta buffer in render_val_text from 16 to 24
bytes so a worst-case ut64 decimal (20 digits + sign + NUL = 22) fits
without tripping -Werror=format-truncation on gcc with -O0.
Also fixes three swapped-argument mu_assert calls in test_pf.c (the
v(N) bitvector tests added in commit 8): the macro signature is
mu_assert(message, test) but those three calls passed (test, message),
which tripped -Werror=format= under -O3 because the int 'test' was
fed to the format string's %s slot.
Adds a second round of regression fixes prompted by a manual audit
against pristine upstream (review feedback after the first pass):
* render_quiet: handle raw byte-sequential types (RZ_PF_UINT128,
RZ_PF_HEXDUMP). scalar_text() has no per-byte case for these so the
default-clause '?' was emitted, making 'pfq Q' and 'pfv Q' print '?'.
render_quiet now mirrors text mode and emits a space-separated hex
stream for raw types.
* Sized pointers (p{2,4,8}): the legacy parser accepted 'p2', 'p4',
'p8' to force a 16/32/64-bit pointer width regardless of ctx.bits;
the new parser only handled bare 'p'. Restored via a new dispatch
branch ahead of bare 'p', recording the override in
RzPfField::bit_width (overloaded -- documented in rz_pf.h). A new
fld_ptr_size() helper consults the override first; read_ptr, the
pointer-deref read path, the STRPTR read path, and
pf_struct_size_impl all use it. 'pf p2p4p8pp2' now consumes
2+4+8+4+2=20 bytes and renders the right value at each position.
* Numeric pointer dereference: '*d4', '*x2', '*u8', etc. (pointer to
fixed-size scalar) only read the pointer itself and never followed
it. The reader now also calls ctx->read_at to fetch the target word
and populate scalars[0]; the text renderer prints both the pointer
literal and the dereferenced value -- '(*0x20) 42' instead of bare
'(*0x20)'. Affects 'pf *d4 ...' and the 'Pointers' / '32 bit twice
then string' tests.
* Pointer-to-struct dereference (*?): pointer-to-struct fields read
the pointer value but stopped there, so 'pf *?' (or typedb forms like
'(troll)Bah' marked with '*') showed only '(*0x30)'. The reader now
recurses through ctx->read_at into a worst-case 4 KiB buffer and
re-invokes read_nested_struct; the renderer prints the nested struct
body after the '(*ptr)' annotation. Recursion is bounded by
ctx->max_depth so cyclic Bah->Bah pointer chains terminate at the
documented limit instead of unbounded recursion. Affects 'nested
struct', 'complex nested struct', and 'flag for nested struct'.
* String pointer (bare 's'): RZ_PF_STRPTR rendering didn't show the
dereferenced target -- output was bare '"hello"' with no pointer
context. Now emits '(*ptr) "string"' (mirroring '*z') when the
deref produced a non-empty body; falls back to bare '""' on
unmapped or empty targets to avoid advertising a phantom pointer.
* render_val_text is_pointer branch: extended to render the nested
struct body for *? pointers and the dereferenced value for numeric
*d4/*x2/*u8 forms.
* Top-level read loop: rz_pf_read passed both 'cur_off + off' (as
read_field's off parameter) AND 'base_addr + cur_off' (as base_addr),
but read_field computes 'val->offset = base_addr + off'. This
double-added cur_off on every iteration past the first, so 'pf 2ic'
reading 5 bytes per iter reported iter 1's nb at offset 0x0a instead
of 0x05, and 'pf 2F' reported iter 1's double at offset 0x10 instead
of 0x08. Data reads were correct; only the displayed/JSON offsets
drifted. The fix: pass base_addr unchanged so val->offset becomes
base_addr + (cur_off + off). Affects 'array obj', 'print n-times a
format', 'pf', 'pf field name', 'JSON output'.
* Bare 'C' (legacy 1-byte unsigned decimal): the legacy parser accepted
'C' as 'print byte as decimal'; the new parser had dropped it. The
'types' test format 'pf fcb1d4C foo bar fool beer plop' was therefore
truncating to 4 fields. Restored as a deprecated alias for 'u1' (with
the standard 'use u1' note).
* pfw dotted-path navigation: the write-mode lookup only matched
top-level field names, so 'pfw gobelin.Buh.first=42' through nested
structs failed with 'field not found'. Replaced with a segment-by-
segment walker that descends through children at each dot. Affects
'write specific element through nested struct'.
* Drop '.pfw' usage from tests. '.pfw' (execute pfw output as rizin
commands) was a legacy convention -- the new pfw writes directly via
rz_io_write_at and emits a human-readable confirmation line, so the
'.' prefix evaluates that confirmation line as a command and fails.
Same direction as the earlier '.pf*' removal. All cmd_pf,
cmd_pf_write, cmd_pf2 tests updated.
* test/db/cmd/cmd_pf 'Register' marked BROKEN with a tracking comment:
the legacy 'r (regname)' looked up CPU registers via
RzPrint::get_register; the new parser repurposes 'r' as raw hex byte
dump and the register-fetch path is gone. Restoring would mean
wiring a new register-lookup hook through RzPfCtx, which is outside
the parser-rewrite scope.
* test/db/cmd/cmd_pf2 'pf F max precision (#13027)' annotated: not a
precision regression. Values differ from the original #13027 expected
output because the legacy 'F' specifier read bytes as LE-then-
reinterpret while the new parser follows the documented 'UPPERCASE
= BE' rule literally; the 17-digit precision -- the actual subject
of #13027 -- is preserved.
EXPECTs updated to match the corrected output across cmd_pf, cmd_pf2,
cmd_pf_write.
* doc/pf.md updated to reflect all DSL and rendering changes:
documents the sized pointers (p/p2/p4/p8), pointer dereference
semantics (numeric, struct, and string), the missing composites
(Q, U, L, n/N), the case-as-endian rule (no standalone endian
directive), the pfw write mode (dotted-path navigation, direct I/O,
no '.pfw'/'.pf*'), and the table of deprecated single-letter
aliases (b, d, o, q, u, i, f, F, w, Z, t, T, X, C).
* librz/type/pf_render.c (new): extracted the render-mode code paths
from pf_parser.c into their own translation unit. Hosts the text /
quiet / JSON / cstruct / DOT renderers, the per-mode helpers
(scalar_text, scalar_json, render_binary, render_guid,
compute_name_width, field_matches, emit_colored*), the RenderCtx
bag, and the public rz_pf_render() + rz_pf_render_json() entry
points. pf_parser.c keeps the parse + read pipeline and shrinks
from 4831 to 3280 lines. Helpers used by both translation units
(is_string_type, is_raw_type, endian_str, pf_vasprintf) move to
pf_parser.h as static inlines.
* librz/type/pf/ (new directory): moved every pf source out of the
librz/type/ top level into a dedicated pf/ subdirectory; only the
legacy format.c stays top-level. meson.build updated with the new
paths and the pf/ include dir.
* Split the monolithic parser into focused sub-grammar TUs, wired
together by the new pf/pf_internal.h (which centralizes the PF_DIAG
diagnostic macro, the shared ReadState, and all cross-TU
declarations):
- pf_parser_string.c string/encoding specs (z / s / Z)
- pf_parser_bitfield.c inline + typed bitfields
- pf_parser_bitvec.c bitvectors v(N)
- pf_parser_array.c array-count resolution
- pf_parser_struct.c nested struct / union reading
pf_parser.c shrinks from 3280 to ~2620 lines and now hosts only the
parse driver, type-spec dispatcher, reader core, context, and the
public utility surface. The TLV TU was migrated off its bespoke
extern/TLV_DIAG block onto the shared header.
* Add an RzStructuredData renderer: pf/pf_render_sd.c with the public
rz_pf_render_sd() (declared in rz_pf.h). (Named _sd, not _sdb: SDB is
Rizin's key/value database, unrelated to RzStructuredData.) It maps the decoded value
vector to the generic key/value document model -- scalars to typed
entries, arrays/bitvectors to arrays, nested structs to sub-maps
with a _type tag, raw/GUID payloads to byte blocks, timestamps to a
formatted string plus a <name>_raw sibling -- so callers get JSON,
YAML, or iterator access for free. Unit and property tests cover it
(see the consistency-pass note below).
* doc/pf.md and librz/type/README.md updated for the new pf/ layout
and the structured-data render mode.
* Consistency pass over librz/type/pf/: deduplicated the structured-data
scalar dispatch (one classifier feeding thin map/array emitters instead
of two parallel ~60-line switches), dropped unused includes
(pf_parser_time.h from the TLV TU, string.h from the struct TU,
rz_endian.h from the SD renderer) and a redundant extern decl now
covered by pf_internal.h, and fixed a dangling-pointer bug in the SD
char path (the classifier returned a pointer into its own by-value
temporary). Expanded coverage: 11 SD unit tests (157 total) and 6 new
theft properties (SD tree non-NULL, JSON structural well-formedness,
top-level object shape, YAML safety, determinism, filter-narrows).
* Split pf_render.c (1548 lines) into per-mode renderer TUs, mirroring
the earlier pf_render_sd.c extraction: pf_render_text.c (text+quiet),
pf_render_json.c (+rz_pf_render_json), pf_render_cstruct.c, and
pf_render_dot.c. A new pf_render.h carries the shared RenderCtx record
plus the cross-TU helper (pf_field_matches, pf_scalar_text,
pf_render_guid) and per-mode entry-point declarations. pf_render.c now
holds only those shared helpers and the rz_pf_render() dispatcher
(~372 lines). meson.build, the file-doc layout listings, and
librz/type/README.md are updated accordingly.
* Second consistency pass: ensured every public RZ_API entry point has a
Doxygen \brief block (added ones for rz_type_format_struct_size and
rz_pf_render_sd) and every file a \file description (added to
pf_parser.h and pf_parser_time.h). Deduplicated the field-size logic:
pf_struct_size_impl's sum and union loops now share one
pf_field_static_size() helper, the enum/bitfield byte-width override is
a single pf_enum_bitfield_width() helper (was inlined 3x), sized-pointer
width routes through the existing fld_ptr_size(), and the inline-bitflag
test shared by the JSON and DOT renderers became the
pf_value_has_inline_bitflags() predicate in pf_render.h. Dropped unused
includes left over from the render split. No behaviour change (157 unit
tests, 146 theft properties, full cmd sweep all green).
* Fix a CodeQL High alert (cpp/integer-multiplication-cast-to-long) in
the scalar-array path selector: the per-element byte delta was computed
as `idx * width` in 32-bit int before being added to the ut64 offset,
so a large array index could overflow int prior to the widening. The
multiplication is now done in ut64 ((ut64)idx * width). No behaviour
change for in-range indices; verified by the path-selection cmd tests.
* Fix a big-endian bug in the structured-data renderer: the bitvector
reader stores each bit in the RzPfScalar union's v_u8 member, but the
SD classifier read it back through v_u64. On little-endian those alias
to the same low byte so it worked by luck; on big-endian (s390x) v_u8
is the high byte of the 64-bit slot, so every set bit read as a large
number and the test_pf_render_sd_bitvec assertion failed. The
classifier now reads v_u8 for RZ_PF_BITVEC, matching the writer; the
other scalar types already read the same width the reader wrote.
* Clean up leftover refactoring-era comments in the renderer TUs: drop
the "Split out of pf_render.c" changelog phrasing from the file-doc
headers (the docs now describe the current layout) and update two
stale references to the pre-rename helper names (field_matches ->
pf_field_matches, render_guid -> pf_render_guid) in comments.
---------
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
Consolidate the Unicode subscript notation used when rendering
bit-vector and float values (the subscript width on a bit-vector
constant, e.g. 0x2c followed by a subscript 8, and the format width
on a float, e.g. .f followed by a subscript 32) into one place, so
the RzIL Unicode exporter, the RzNum value printer, and the
RzNum->RzIL lift cannot drift apart.
RzUtil gains the single source of truth:
* rz_str_append_subscript() / rz_str_append_superscript() /
rz_str_subscript() render a number as Unicode subscript or
superscript digits;
* rz_bv_width_subscript() / rz_bv_as_unicode_string() build a
bit-vector's width subscript on top of the str helper;
* rz_float_format_subscript() renders a float format's width
subscript (16/32/64/80/128, with the decimal-format marker),
reusing the same digit renderer.
The RzIL Unicode exporter (il_export_string_unicode.c) is switched
fully onto these: every append_subscript() call site (bit-vector
constant width, cast length, memory indices) now goes through
rz_str_append_subscript(), and the hardcoded per-format subscript
macro is replaced by rz_float_format_subscript(). The exporter's
private subscript-digit table and ut32 glyph helper are removed, so
there is no longer a second, parallel implementation to keep in sync.
Signed-off-by: Anton Kochkov <anton.kochkov@gmail.com>
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
The existing integration tests only reached the per-type graph
builders through the rz_core_graph() dispatcher. Add tests that call
rz_core_graph_callgraph/datarefs/coderefs/importxrefs directly and
assert they agree with the dispatcher, plus coverage for the
rz_core_graph_to_dot_str()/rz_core_graph_to_sdb_str() serializers.
Completes the work on rizinorg/rizin#992.
The visual bit editor (V d 1 / V b1) ignored the active color theme,
scr.utf8 and cfg.bigendian, always showed ESIL even when RzIL was
available, overflowed the IL onto a single line, and drew the cursor
position as a detached underline. This reworks it to match the styling
of the commands users already know and adds a cursor-position summary.
- Byte order on screen follows cfg.bigendian: big-endian shows the
byte at the current offset on the left, little-endian reorders so the
MSB is on the left. Cursor movement and every byte-write key go
through a MEM_BYTE() mapping so edits land on the byte under the
cursor regardless of endianness; the buffer stays in memory order.
- chr/dec/hex byte cells are colored by value via rz_print_byte_color,
the same helper px uses. The asm line keeps rz_asm_colorize_asm_str.
- The rzil line is colored like plf: core_colorify_il_statement is
split into an address-prefix wrapper and a body-only colorizer
(rz_core_il_colorize_body) shared by both. Long IL is soft-wrapped at
the terminal width, breaking between top-level seq children and
indenting continuations.
- RzIL is preferred over ESIL; the analysis op now also requests
RZ_ANALYSIS_OP_MASK_IL. When neither is available the line is omitted.
- The bit under the cursor is reverse-video highlighted (a graphics
attribute emitted regardless of scr.color); padding bits are dimmed
via pal.comment. The empty-bit padding glyph is the middle-dot when
color or Unicode is on, else a plain period.
- The position indicator now sits between the two bit rows as a single
marker pointing at both, and a compact info line summarizes the
cursor: "byte N - nibble H|L - bit B [P] - LE|BE".
- The ? help is expanded with a description of the mode and a legend
for the position line; the key table is colored like the other
visual help screens via rz_core_visual_append_help.
- Fix a leak of the asm op on the q/Q exit path.
Closes#6361
Adds full-function disassembly (pdf) tests for non-trivial functions on
both CPUs, to lock in correct rendering of basic-block boundaries,
reflines (forward and backward branches), call/data/string cross-
references and operand formatting.
C55x (emulateme_nostd.ccsv5.c55x.ticoff2.dbg.coff):
- sym._uart_write_hex: a compact two-block function with a forward
conditional branch, an unconditional branch refline and a call ref.
- sym.___text: a four-block loop with conditional branches both ways,
a back-edge, a call into another function, data references and
parallel-instruction (||) operands.
C55x+ (coff2/19_emulateme_nostd.obj):
- sym._c_strlen: a four-block scan loop with a back-edge and forward
and backward bcc reflines.
- sym._main: a five-block function with nested conditional branches
and two call references (_c_strlen, _decrypt).
All four render cleanly with the existing decoder and analyzers; no
behavior change was needed.
Gives the C55x / C55x+ disassembler a Capstone-style per-instruction
identifier and threads it through to analysis, so the analyzers dispatch
on named operations instead of raw opcode bytes.
What lands
==========
* librz/arch/isa/tms320/tms320c55x_insn.{c,h} -- a shared TMS320C55InsID enum
(TMS320C55_INS_AADD, TMS320C55_INS_B, TMS320C55_INS_CALL, ...) covering
C55x and C55x+, plus tms320c55x_insn_name(), tms320c55x_insn_id_from_syntax()
and tms320c55x_insn_optype(). The prefix is TMS320C55_ / tms320c55x_
rather than TMS320_ / tms320_ because the other TMS320 families (C54x,
C64x, C28x) have substantially different instruction sets; the enum,
the file and these helpers are C55x/C55x+ specific.
* insn_head_t gains an .id field; every head entry in c55x/table.h is
tagged with its TMS320C55InsID. The disassembler resolves the decoded
instruction ID from the emitted mnemonic (so multi-form leading bytes
such as 0x95 -> intr/trap, 0x48 -> ret/reti/rpt, 0x50 -> sftl/popboth
resolve to the exact instruction, which a static byte->head map cannot).
tms320_dasm_t gains an insn_id field and tms320_dasm_insn_id() accessor.
* tms320_c55x_insn_id_decode() / tms320_c55x_plus_insn_id_decode() let the
analyzers resolve the named ID for a byte sequence via a cached decoder
instance.
* The C55x analyzer's dispatch is rewritten from switch(opcode_byte) to
switch(TMS320C55InsID). Both analyzers set op->id to the named ID, the
same way the Capstone-based plugins (e.g. c64x) populate op->id.
* The C55x+ analyzer keeps its byte-level dispatch for control flow,
stack deltas and operands, but takes the final arithmetic/logical/move/
multiply/stack type from tms320c55x_insn_optype(op->id). A leading byte on
C55x+ encodes several instructions (selected by operand bits), so the
byte switch alone cannot tell ADD from SUB, AND from OR/XOR, or AMOV
from ASUB; the decoded id can.
Bugs fixed
==========
Driving dispatch / typing from the decoded id fixes a number of latent
mis-classifications the raw-byte switch had masked, verified against the
TI dis55 reference disassembler and the C55x+ documentation:
- 0x50 0x66 (psh Tx) was typed SHL; now a stack push (corrects
sym._main's computed stackframe in the rel.stripped.coff test).
- 0x48 0x05 (reti) was typed REP; now RET.
- 0x95 0x0F / 0x8F (intr vs trap) distinguished by the decoder.
- C55x+ 0x7B: byte1 bit7 selects LD (mov) vs add/sub, and within
add/sub byte2 bit7 selects sub; was always typed add/mov by nibble.
(TI SWPU104 Table 7-2, opcode 01111011.)
- C55x+ 0xD2 mar(XDAa op k24): address-register modify, now LEA like
AADD / AMOV; was typed SUB. (Table 7-2, opcode 11010010.)
- 57 further C55x+ arith/logical/move/stack contradictions found by
cross-referencing a Motorola Droid (Wrigley3G) C55x+ baseband dump
against the decoder are resolved by the optype override.
* DELAY is a memory-delay MOVE per TI SWPU104 sec.6.7.1 (Memory Delay,
grouped under Move Operations): it copies Smem to Smem+1. Both
analyzers now mark it as a memory access (width 2, write) and drop the
spurious FAMILY_CPU (CPU is already the default family).
Tests
=====
New checks assert op->id carries the right TMS320C55_INS_* value on each
CPU, plus regression tests for every byte/decoded-id mismatch fixed above
(c55x: 0x50/0x48/0x95/0xb6; c55x+: 0x7b/0xd2). The c55x opcode-
classification and stackframe expectations are updated to the corrected
output.
Cross-reference
===============
TI SPRU374 'TMS320C55x DSP Mnemonic Instruction Set Reference Guide'
TI SWPU086 'TMS320C55x+ DSP Algebraic Instruction Set Reference Guide'
TI SWPU104 'TMS320C55x+ DSP Mnemonic Instruction Set Reference Guide'
Short summary of the three cpu= values supported by the tms320 arch
plugin (c55x, c55x+, c64x), the silicon they map to, and a note on
the two distinct meanings of 'c55x+': TI's publicly-named C55x DSP
Core+ shipping in current low-power C55x parts (C5504-C5545), and
the pre-release Ryujin revision documented in TI SWPU086 / SWPU104
that lives in older modem silicon. The rizin plugin handles both.
TI's cl55 compiler writes COFF files with little-endian file and
section headers, but the data inside DWARF sections (.debug_info,
.debug_abbrev, .debug_frame, .debug_line) is laid out as 16-bit
big-endian words. The DWARF parser was reading the unit length as
LE and getting absurd values, then bailing out with zero
compilation units.
Result: afvl, avgl, afs all returned empty for every
emulateme*.ticoff2.dbg.coff in rizin-testbins even though the
file had a fully-populated .debug_info section, the abbrev
tables parsed cleanly, and the per-architecture DWARF register
mapping (added separately) was available.
Override bf_bigendian() for arch=tms320 so the DWARF endian reader
swaps correctly. The COFF info struct stays in the same
little-endian state it always was; only the dwarf reader's view
flips.
After this fix, cl55-compiled TI COFF v2 binaries with debug info
yield:
afvl @ dbg.main -> arg int argc @ ar4
afs @ dbg.main -> int main(int argc);
avgl -> 7 globals (uart_address, seckrit, _lock, ...)
pdf @ <fn> -> renders with signatures, named args, reflines,
resolved call targets
Rewrites the C55x and C55x+ analysis classifiers as pure byte-level
dispatch -- no mnemonic-string matching, no round-trip through the
disassembler -- and adds the supporting infrastructure they need to
produce useful RzAnalysisOp metadata.
What lands
==========
* librz/arch/isa/tms320/c55x/c55x_analysis.{c,h} -- C55x baseline
classifier, ~360 lines, 256-entry size table extracted from the
decoder's table.h.
* librz/arch/isa/tms320/c55x_plus/c55plus_analysis.c -- C55x+
classifier rewritten in the same shape, ~470 lines covering 90+
opcodes with byte-level disambiguation for 0x02 / 0x03 / 0x74 /
0x76 / 0x7B / 0xC5.
* librz/arch/isa/tms320/tms320_dwarf_regnum_table.h plus a hook in
librz/arch/dwarf_process.c -- TI cgt55 ABI DWARF register-number
mapping, so the cl55 compiler's .debug_info variable locations
resolve into rizin register names instead of returning the dummy
"?" placeholder.
* librz/arch/p/analysis/analysis_tms320.c -- thin dispatcher that
picks the per-cpu classifier and stops carrying the tms320_dasm_t
engine in analysis state.
Why a byte-driven classifier
============================
The old classifier round-tripped through the disassembler and did
strncasecmp() on the mnemonic string. Three problems:
1. It kept a tms320_dasm_t engine alive in the analysis context
just to read its 'syntax' buffer after every classify call.
Removing it shrinks the per-analysis state and removes a
tms320_dasm_init/_fini pair from the analysis_init/_fini path.
2. It only set op->type -- never op->jump, op->fail, op->stackop,
op->stackptr, op->val, op->eob. Basic-block formation followed
only the most obvious control flow, and call/ret/push/pop
semantics were invisible to higher-level analysis.
3. It couldn't disambiguate predicated versus unconditional calls:
the disassembler emits 'callcc' vs 'call', but the substring
match missed the conditional fail-path for CALLCC.
The new classifiers fix all three:
- Read the leading byte (and second-byte refinements where the
encoding family is shared) directly from buf.
- Resolve jump and call targets from BE-stored displacement and
absolute fields, with correct sign extension for the 8-bit and
16-bit relative forms.
- Read 24-bit absolute targets via rz_read_at_be24().
- Set op->fail = addr + size for every conditional jump/call,
op->eob = true for unconditional branches and RET so basic-block
walkers terminate correctly.
- Track the stack: PSH/POP per ISA cluster, CALL/CALLCC +2,
RET/RETI -2.
- Capture INTR/TRAP immediates in op->val via set_imm().
- Disambiguate sub-opcodes that share a leading byte by reading
the relevant bits of the second byte. For C55x, the most
notable case is 0x48 (RPT/RPTADD/RPTSUB/RET/RETI) which uses
bits 0-2 of byte 1; for C55x+ the disambiguations are 0x02,
0x03, 0x74, 0x76, 0x7B and 0xC5.
- Handle parallel-prefix bytes (odd-valued leading bytes below
0x80 in C55x like 0x03, 0x05, 0x07, 0x11, ...) by treating
them as a 1-byte prefix and dispatching on byte 1 so paired
'|| retcc', '|| bcc', etc. classify correctly.
Both classifiers ship analyzer helpers (set_cjmp, set_call, set_jmp,
set_ret, set_cret, set_push, set_pop, set_imm, set_mem_width,
set_dst_reg, set_ireg, set_dir, set_disp) so each opcode entry fills
the RzAnalysisOp ptr / val / stackop / stackptr / fail / eob fields
uniformly across both architectures.
DWARF register mapping
======================
Loading any cl55-compiled TI COFF v2 with debug info (every
emulateme*.ticoff2.dbg.coff in rizin-testbins) used to fire:
ERROR: No DWARF register mapping function defined for tms320 32 bits
per variable, because dwarf_process.c had no entry for arch=tms320.
The new tms320_dwarf_regnum_table.h covers the cgt55 ABI numbering:
AC0-AC3, T0-T3, AR0-AR7, SP/SSP/CDP, BK03/BK47/BKC, DP/PDP, CSR,
BRC0/BRC1, TRN0/TRN1, RPTC, IER0/IER1, IFR0/IFR1, DBIER0/DBIER1,
IVPD/IVPH, ST0_55..ST3_55 (42 entries). Reach into the table is
guarded; out-of-range numbers fall back to NULL so the caller
surfaces the dummy "?" instead of confidently picking the wrong
register.
Wrigley3G coverage
==================
Validation against a 3.1 MB Wrigley3G baseband firmware (Motorola
Droid A855, MSG39UPEU_A1.19_1.80, partition CG45.img) found 31
leading-byte values producing real instructions classified as NULL.
The c55x+ classifier here covers those:
0x50-0x5F MOV memory/register cluster
0x88, 0x8A MOV ACx <-> mem high/low halves
0x8C ADD with carry, mem -> ACx
0x97 Dual-memory MOV (parallel)
0xA0 MOV with parallel dual addressing
0xAC, 0xAD MOV #k16, ACx (long immediate)
0xB4, 0xB5 MOV with rounding and shift
0xB6, 0xB7 ADD with shift (T-register or immediate)
0xC0, 0xC2, 0xC4 ADD #k16 with shift slots
0xCC Packed ADD :: MOV dual-instruction encoding
0xD0 MOV ACx, dbl(*(#abs24))
0x2E, 0x2F XCCPART predicated execute
0x0B, 0x23 Wrigley silicon pseudo-ops (TRAP)
0xC6 BFXTR / BFXPA bit-field extract (MOV)
The 0x03 family classifier extends from a 4-bit (0xF0) to a 6-bit
(0xC0) mask so the full encoded range resolves:
0x03 0x00-0x3F INTR #k5
0x03 0x40-0x7F TRAP #k5
0x03 0x80-0xBF SWAP register pairs
0x03 0xC0-0xFF SIM_TRIG (Wrigley-specific simulator trigger)
Coverage on Wrigley3G rises from 94.4% to 97.4% (2000-sample
random survey).
Tests
=====
Two new test suites land alongside the classifiers:
test/db/analysis/tms320.c55x_32 11 tests (batched)
test/db/analysis/tms320.c55x+_32 13 tests (batched + binary
fixtures)
Tests are intentionally batched -- each test bundles 10-12 opcode
checks behind one rizin process spawn instead of one per check.
That brings both suites down to under 0.5 seconds combined.
The c55x+ suite includes six binary-fixture tests against the
companion rizin-testbins drop-in tms320/coff2/*.obj corpus,
covering function discovery (afl), stack-pointer tracking
(afvs / afS), data-section walk (iS), and globals enumeration
(is). The c55x suite covers tms320/emulateme_nostd.ccsv5.c55x
.ticoff2.dbg.coff from the existing rizin-testbins tree.
Cross-reference
===============
TI SPRU374 'TMS320C55x DSP Mnemonic Instruction Set Reference
Guide' (publicly available) -- C55x baseline.
TI SWPU086 'TMS320C55x+ DSP Algebraic Instruction Set Reference
Guide' (May 2005) -- C55x+ instruction encodings.
TI SWPU104 'TMS320C55x+ DSP Mnemonic Instruction Set Reference
Guide' (December 2006) -- C55x+ mnemonic forms.
The c55x_plus disassembler glue layer (c55plus.c) is rewritten as a
thin shim over the th0rpe c55plus_decode() walker:
- Drop the ad-hoc ctype tolower() loop; use rz_str_case() to
lower-case the walker's mixed-case mnemonics in one call.
- Reorder the includes; drop the unused ones.
- Scope local variables to where they are actually used.
- Move the global setup (ins_buff / ins_buff_len) into the same
block as the c55plus_decode() call so the dataflow is obvious.
While here, fix a pre-existing off-by-one in utils.c. The hex-digit
lookup table was declared as a 17-character string with a duplicated
leading '0':
static char hex_str[] = "01234567890abcdef";
This shifted every nibble >= 0xA by one position in the table:
hex_str[10] = '0' (should be 'a')
hex_str[11] = 'a' (should be 'b')
...
hex_str[15] = 'e' (should be 'f')
hex_str[16] = 'f' (unreachable)
Result: get_hex_str(0xff) returned "ee", get_hex_str(0xab) returned
"0a", and every disassembled '.byte 0xNN' for an unknown opcode with
nibbles >= A came out wrong. The function is used from
c55plus_decode.c on the unknown-opcode fallback path
(hash_code == 0x223), so the bug surfaces whenever the decoder bails
out and emits a raw byte.
Fix: use the correct 16-character table "0123456789abcdef" and rename
the static to hex_digits to make the role obvious. While here, tidy
strcat_dup() to use bitwise tests on the n_free bitmask so the
'3 = free both' contract is enforced uniformly, add docstrings to both
helpers, and drop the redundant memcpy length guard that was a no-op
for non-NULL length-zero strings.
No behavioural change for either piece beyond the bug fix above.
The c55x_plus decoder shipped with two private string helpers in
utils.c:
strcat_dup(s1, s2, n_free) - allocate s1+s2 and optionally free
inputs, with a bitmask controlling
which of s1/s2 are released
get_hex_str(n) - format the low 8 bits of n as a
two-character lowercase hex string
Both have direct equivalents in rz_util:
strcat_dup(s, lit, 1) -> rz_str_append(s, lit)
strcat_dup(lit, s, 2) -> rz_str_prepend(s, lit)
strcat_dup(s1, s2, 3) -> rz_str_append_owned(s1, s2)
strcat_dup(s1, s2, 1) where
s2 is also owned and freed
manually right afterwards -> rz_str_append_owned(s1, s2)
get_hex_str(n) -> rz_str_newf("%02x", n & 0xff)
This commit converts all 56 strcat_dup call sites in
c55plus_decode.c and decode_funcs.c plus the single get_hex_str
site, then deletes utils.c and utils.h entirely.
While here, replace several local sprintf-into-stack-buffer +
rz_str_dup patterns with direct rz_str_newf calls:
- get_AR_regs_class1: was malloc(50) + sprintf per case, now a
single rz_str_newf per case returning the result directly.
The function is reduced from 34 lines to 14.
- get_AR_regs_class2: same pattern, reduced from 130 lines to
79 with no allocation needed at the top.
- get_token_decoded case 40/48, 70/72/80, 41/73: sprintf into
a 512-byte stack buffer then rz_str_dup -> single rz_str_newf.
- decode_funcs.c case 2 of get_status_regs_and_bits: was
calloc(50) + sprintf, now rz_str_newf.
The 512-byte stack buffer 'buff_aux' in get_token_decoded becomes
unused and is removed.
C55PLUS_DEBUG, the only useful symbol that used to live in utils.h,
moves to ins.h (which all c55x_plus translation units transitively
include). ins.h gains a direct <rz_util.h> include so the rest of
the headers don't need to pull it in indirectly.
No behavioural change. Full c55x_plus test regression passes:
asm/tms320_c55x+_32 104/104, analysis/tms320.c55x+_32 45/45,
parity against TI dis55.exe v4.3.6 on the 13-source testbins corpus
remains 140/140.
COFF C_EXT (external/global) symbols carry their function-or-not
status in two places: the DTYPE field of n_type (the ISFCN bit), and
implicitly by the section they live in. The existing code only
honoured the DTYPE bit:
ptr->type = DTYPE_IS_FUNCTION(s->n_type) || !strcmp(name, "main")
? RZ_BIN_TYPE_FUNC_STR : RZ_BIN_TYPE_UNKNOWN_STR;
C and C++ compilers set the ISFCN bit when emitting object code, so
this works for compiled output. But assembler-emitted globals -- TI's
asm55p / cl55, the GNU GAS COFF backend on legacy targets, and any
hand-written .s -- often leave n_type at zero. Every assembly symbol
then comes back as RZ_BIN_TYPE_UNKNOWN_STR, and 'aaa' has no way to
distinguish a function from a data label. Concrete example: with TI
asm55p output for the C55x+ test corpus, none of the eight .global
labels in 03_branches_calls.s (_short_ret, _short_branch, _short_call,
...) were promoted to functions, so analysis only found those
reachable by following control flow from a hard-coded entry.
Add a section-flag fallback: if the symbol's containing section has
COFF_SCN_CNT_CODE set (i.e. it's a .text / code section), the symbol
is a function. This keeps the DTYPE check as the primary signal but
fills the gap on assembler output.
Verified with TI C55x+ .obj files: all .global labels now appear as
RZ_BIN_TYPE_FUNC_STR and are picked up by 'aaa'. No regression on
the standard COFF test suite (Windows i386/amd64 .obj from compiled
C output).
The COFF section table walker assumed every COFF section header is
40 bytes -- the standard COFF1 form. The TI Common Object File Format
(SPRAAO8) extends this for TI COFF v2 (file magic 0x00C2) used by the
TI tools (asm55, cl55, cl6x, cl28, etc.):
------------------------------------------------------------------
offset size field standard COFF1 TI COFF v2
------------------------------------------------------------------
0x00 8 s_name char[8] char[8]
0x08 4 s_paddr ut32 ut32
0x0c 4 s_vaddr ut32 ut32
0x10 4 s_size ut32 ut32
0x14 4 s_scnptr ut32 ut32
0x18 4 s_relptr ut32 ut32
0x1c 4 s_lnnoptr ut32 ut32
0x20 2/4 s_nreloc ut16 ut32 <-- widened
0x22 2/4 s_nlnno ut16 ut32 <-- widened
0x24 4 s_flags ut32 ut32
0x28 - reserved - ut16 <-- new
0x2a - mempage - ut16 <-- new
------------------------------------------------------------------
TOTAL 40 48
Reading TI COFF v2 with the standard 40-byte stride walks the section
table off-by-8 per section, producing nonsense values for every
section after the first. In practice, vaddr and size for the second
section spill into the next section's name field, surfacing as
attention-grabbing decimal values like 0x7461642e (ASCII '.dat'
reversed -- bytes from the upcoming '.data' name being interpreted as
a size).
Concrete reproducer with TI asm55p output:
$ wine asm55p.exe -v5505 03_branches_calls.s
$ rz-bin -S 03_branches_calls.obj # before this fix
...
0x00000061 0x7461642e 0x00000040 0x7461642e ... # garbage size
...
$ rz-bin -S 03_branches_calls.obj # after this fix
0x000000fa 0x34 0x00000030 0x34 -r-x .text
...
Fix: add a parallel coff_init_scn_hdr_ti() that reads the 48-byte
layout, with nreloc/nlnno as ut32 (clamped to ut16 because the rest
of the COFF code keeps them as 16-bit and counts above 64K are not
encountered in practice). bin_coff_init_scn_hdr() picks the variant
based on coff_is_ti_machine() -- the same predicate already used in
the file-header parser to consume TI's f_target_id field.
Tested with TI asm55p output for C55x, C55x+, and C5500 (machine ids
0x9c, 0xa1, and TI_1/TI_2 file magics) -- sections now parse with the
correct sizes, vaddrs, and flags, and the resulting binaries open
cleanly under 'rizin -A' for analysis.
The c64x branch of tms320_reg_profile() declared '=PC pc' but only
ever registered pce1 -- pc was never one of c64x's declared registers.
This raised:
WARNING: Invalid alias given in register profile: pc.
on every rz-asm / rizin invocation of any tms320 cpu, including c55x
and c55x+, because the warning fires during analysis init before the
cpu-specific branch of the profile is selected.
The C64x ISA uses pce1 (Program Counter Extension 1, .32 at index
545) as its program counter -- the only PC-named control register
declared in the c64x profile -- so '=PC pce1' is the correct alias.
The companion test/db/analysis/tms320.c64x_32 'arp' expectation is
updated to match.
The c55x / c55x+ branch (is_c5000) is unchanged -- those profiles do
declare 'pc' as a 24-bit program counter, so '=PC pc' stays valid
there.
This change is independent of the c55x_plus work in the rest of this
series; it would be a useful cleanup on its own. Included here because
it surfaces immediately whenever any of the new c55x+ analysis tests
are run.
Detected with valgrind, some element assignments could memcpy with
identical addresses. This is usually a no-op in practice, but
theoretically undefined behavior.
* Add option to replace (geometric) invalid values with another value.
* Add geometric Mean and Std Dev to benchmarks.
* Add Doxygen documentation
* Add a README to the bench dir as intro.
empty_cond was never signalled and the termination depended only on
the timeout in rz_th_queue_close_when_empty() causing a re-check of
emptiness.
However, the rz_th_cond_timed_wait() implementation, which was used
there, was flawed because it expected a relative timeout but passed that
directly to pthread_cond_timedwait() which expected an absolute time
value, practically causing it to time out immediately. Depending on the
pthread_cond implementation, this possibly created a situation where the
mutex could never be acquired by another thread, effectively causing a
deadlock. This behavior was observed on Mac OS X 10.5 (ppc) when running
the test_core_bin test.
We solve this by not using a timeout at all and signalling the condition
variable for all waiting threads at the appropriate time.