* librz/core,cmd: add pf-aware autocompletion arg types
The `pf` family of commands (`pf`, `pf-`, `pfa`, `pfc`, `pfd`, `pf.`,
`pfn`, `pfo`, `pfs`, `pfv`, `pfw`) takes either a registered named
format, a `<format>.<field>[<idx>]...` path, or a Format Definition
File. None of these were autocompletable: the arg type was always
`RZ_CMD_ARG_TYPE_STRING`, so tab on `pf. <TAB>` did nothing and the
user had to remember every format and field name by hand.
Add three new arg types and wire them to the cmd_descs:
* `RZ_CMD_ARG_TYPE_PF_FORMAT_NAME` enumerates the named formats
from `rz_type_db_format_all()` and filters by the partial prefix.
Used by `pf-`, `pfa`, `pfc`, `pfd`, `pfn`, `pfs`, `pfv`.
* `RZ_CMD_ARG_TYPE_PF_FORMAT_PATH` is the path-aware completer for
`pf.` and `pfw`. Both accept arbitrarily deep paths of the form
`name[.field[<idx>]?]*`, walking through nested struct fields
(e.g. `pf. troll.str[1].two` follows the same path syntax that
`pf_path_navigate` accepts at runtime). The completer:
- Finds the last `.` in the partial input; everything before
it (inclusive) is the committed path, what follows is the
segment being completed.
- Walks each committed `name[N]?` segment in turn, looking up
STRUCT fields' `type_name` in the typedb and re-parsing the
referenced format. Descent past a scalar or an inline struct
(no `type_name`) returns no options.
- Offers the resolved format's field names for the tail.
- Returns nothing when the tail contains `[` or `]` -- the
user is mid-index or mid-descent, where identifier
completion would produce a syntax error if accepted.
- Suppresses the trailing space after each successful
completion so `.` can be typed next without an inserted
space getting in the way of descent.
- Rewrites `res->start` past the last `.` so completion only
replaces the tail; the committed path stays put.
* `RZ_CMD_ARG_TYPE_PF_FDF_FILE` lists the basenames of FDFs found
in the user's home formats dir and the system formats dir,
mirroring the search order used by `pfo` itself; files present
in both locations are reported once via the same `HtSU` de-dup
used in `cmd_print_format_file_handler`. Used by `pfo`.
All three new types are added to `CD_ARG_LAST_TYPES` in
cmd_descs_util.py. The generator sets `RZ_CMD_ARG_FLAG_LAST` on the
final arg of a command whenever that arg's type is in this set;
that flag tells the runtime arg-preprocessor to merge any trailing
whitespace-separated tokens into a single argv slot. This is the
same implicit-FLAG_LAST treatment `RZ_CMD_ARG_TYPE_STRING` already
gets, and it is what keeps invocations like `pfc zd4x8 foo bar cow`
(five tokens, one logical format-with-names argument) working --
without it the cmd parser would reject the extra tokens with "Wrong
number of arguments". The implicit merge does not interfere with
autocompletion: the completer still receives the partial input up
to the cursor and prefix-matches against it, and the format-path
completer's dot/bracket scan is unaffected by whitespace.
The path completer uses a small helper, `pf_path_seg_consume`, that
parses one `name[N]?` segment with safe handling of unterminated `[`,
empty `[]`, non-numeric indices, and end-of-input. `pf_resolve_path_format`
walks all committed segments and returns the RzPfFormat the caller
should complete against; the resolver is the same shape as the
`pf_path_navigate` walker in librz/type/pf/pf_parser.c, except it
operates on RzPfFormat trees (typedb names) rather than RzPfValue
trees (decoded data), so it can run before any read has happened.
The completers live in `cautocmpl.c` next to the existing type-name
completers (`autocmplt_cmd_arg_struct_type` and friends) and follow
the same loop+strncmp pattern. The dispatcher entry for
`RZ_CMD_ARG_TYPE_FOLDER` was missing an explicit `break;` and would
fall through to `default`; harmless, but fixed in passing so the
three new cases sit cleanly above `default:`.
Note on dot-search direction: `rz_sub_str_rchr` is `start..end`
range search returning the FIRST hit, not a right-to-left "find
last" -- the `r` is for "range", not "right". The path completer
needs the last dot, so it walks the buffer backwards itself.
* test/integration: cover pf autocompletion
Thirteen new tests in test_autocmplt.c exercise the three new pf arg
types, including the multi-segment path resolver:
Format name completion:
* `pf_format_name` -- `pfn ut_<TAB>` after registering two formats
confirms both are offered.
Single-segment path completion:
* `pf_format_path` -- three-phase walk through `pf. ut_path<TAB>`,
`pf. ut_path.<TAB>`, `pf. ut_path.cou<TAB>`, covering name-only,
dot-only, and dot-with-prefix. Verifies that `res->end_string`
is empty in the name phase (so `.` can be typed next without an
inserted space) and that `res->start` advances past the dot in
the field phase (so the completion only replaces the field
portion).
* `pfw_format_path` -- the same `<format>.<field>` syntax must
work on the write side too; confirms the PATH completer fires
for `pfw` and is not pf.-specific.
* `pf_format_path_empty` -- bare `pf. <TAB>` lists every
registered format. Snapshots the baseline count first so the
assertion stays robust against any default formats the type DB
might seed.
* `pf_format_path_unknown_name` -- `pf. nonexistent.<TAB>`
returns an empty option list rather than crashing or leaking
diagnostics.
* `pf_format_path_anon_field` -- formats whose fields don't all
have names (e.g. a `.` skip slot) must be iterated safely; the
named fields are offered and the anonymous slot is silently
dropped.
Multi-segment / nested-struct path completion:
* `pf_format_path_nested` -- two-level descent through a STRUCT
field whose `type_name` references another registered format,
parsing the child format and offering its fields.
* `pf_format_path_three_levels` -- three-level descent narrows
correctly: A -> B -> C, then filter C's fields by a prefix.
* `pf_format_path_array_index` -- `pf. troll.str[1].<TAB>` mirrors
the existing cmd_pf2 runtime test; the array index in the
middle segment is parsed and skipped (it doesn't change the
target type).
* `pf_format_path_inside_brackets` -- cursor inside an
unclosed `[` returns no options (mid-index).
* `pf_format_path_after_close_bracket` -- cursor right after `]`
without a trailing `.` also returns no options (mid-descent).
* `pf_format_path_through_scalar` -- descent past a scalar field
is meaningless and returns no options.
The tests use plain `rz_core_new()` (the real cmd_descs already
registers all `pf*` commands), matching the pattern used by
`test_autocmplt_eco_themes`. Format strings use the parser's
"specifier-then-names" form (no internal whitespace in the spec
region) so `rz_pf_parse` produces the expected field count.
* doc,librz/core: align pf docs and `pf?` help with the parser
The standalone reference doc/pf.md and the in-tree `pf?` help (driven
by the details: block in librz/core/cmd_descs/cmd_print.yaml) had
drifted from each other and from what the parser actually accepts.
Both are now consistent with librz/type/pf/pf_parser.c.
Specific corrections:
* `n` family. doc/pf.md claimed `N1`/`N2`/`N4`/`N8` existed as BE
counterparts to `n1`-`n8`, and that bare `n`/`N` defaulted to
`ctx.bits/8`. The parser handles only `n{1,2,4,8}`; all four
forms are context-endian (follow `ctx->big_endian`), and bare
`n` produces "unknown specifier". Rewrite the section to match,
and explain why context-endian is the right choice for header
readers like ELF.
* Deprecation list. The `pf?` "deprecation" note listed `c, s, z`
among the deprecated bare-letter codes -- they are not deprecated
(`c` is the current 1-byte-as-char specifier, `s` is the current
pointer-to-zstring, `z` is the current inline zstring). It was
missing `C, i, Z, X, F, T` which the parser does warn on.
doc/pf.md had `c` in its table marked "unchanged" (so it was
visibly inconsistent with itself) and was missing the `x` row.
Both lists now mirror the parser's PF_DIAG(DEPRECATED) call
sites: b, C, d, f, F, i, o, q, t, T, w, x, X, Z.
* TLV `h=`. `pf?` said "h=v/a (header inclusion)" (two options);
the parser accepts `v` (value only, default), `l` (length covers
len+value), and `a` (length covers tag+len+value). doc/pf.md
already listed all three; help now matches.
* `v(N)` bitvector. doc/pf.md documented this in detail (the
1..4096-bit-wide field type used for things like ELF
`DT_FLAGS_1`, PE characteristics, page-allocation maps), but the
`pf?` help didn't mention it at all. Added to the DSL extensions
section.
* Pointer widths. doc/pf.md documented `p2`/`p4`/`p8` explicitly;
`pf?` only mentioned bare `p` with a "size from ctx.bits"
parenthetical. Help now lists the four forms in one entry.
* GUID layouts. Both docs claimed `G(le)` was "all little-endian",
but the renderer treats `G(le)` and `G(ms)` identically -- D4
(the trailing 8 bytes) is always in buffer order regardless of
layout. Document the actual behaviour rather than the implied
one; this is a description fix, not a code change. Anyone who
wants the byte-reversed-D4 reading can still file it as a
follow-up bug against the renderer in pf_render.c.
cmd_descs.[ch] is regenerated automatically by the custom_target rule
when cmd_print.yaml changes; the .c diff in this commit is the result
of that regeneration (5 lines of comment text inside the existing
detail entries).
format.c can already turn an RzType into a pf format string. Add the
inverse in the same file, next to its counterparts:
rz_type_format_to_c_declaration() parses a pf format string with
rz_pf_parse() and emits an equivalent C struct/union declaration built
from the standard fixed-width types (uint8_t, int32_t, float, ...).
The conversion is structural: it consumes only the parsed RzPfFormat
(field kinds, widths, array counts, pointer flags), never a byte buffer,
so it runs without any target data. Every field kind the engine produces
is mapped; a handful of specifiers with no exact static C form are mapped
best-effort (documented inline): @N alignment is dropped, an
unknown-length inline z string becomes char *, LEB128 widens to its
largest decoded integer, and ?(Name) / E(Name) emit struct/enum
references that must themselves be defined to parse.
Added new 'tdf' command that exposes that conversion to the user
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* librz/type: rewrite pf format parser
Replace the legacy print-format engine in librz/type/format.c with a
clean three-stage pipeline (parse -> read -> render) and a properly
typed DSL.
The new pipeline:
rz_pf_parse(str) source string -> RzPfFormat
rz_pf_read(fmt, buf, len, ctx) RzPfFormat -> RzPfValue[]
rz_pf_render(vals, mode, opts) RzPfValue[] -> printable string
rz_pf_format() one-shot for all of the above
Public API lives in <rz_pf.h> (pulled in transitively via <rz_type.h>).
DSL improvements over the legacy scheme:
* Sized integers with explicit endianness via case:
x4 d4 u4 o4 b4 (LE) vs X4 D4 U4 O4 B4 (BE).
* Sized floats: f2 / f4 / f8 (and BE F2 / F4 / F8).
* Encoding-aware strings: z(utf8) / z(utf16le) / z(utf16be) /
z(utf32le) / z(utf32be) / z(latin1) / z(ebcdic).
* Length-prefixed strings: z[N] reads a N-byte LE prefix then body.
* Parameterized timestamps: t(unix32) / t(unix64) / t(unixms) /
t(unixus) / t(unixns) / t(filetime) / t(dos) / t(hfs) /
t(oletime) / t(webkit) / t(cocoa). ntfs alias maps to filetime.
* GUID: G(ms) (Microsoft mixed-endian, default), G(be), G(le).
* Inline bitfields: B4(FLAG_A=1,FLAG_B=2,FLAG_C=4).
* Alignment: @N advances to next multiple of N.
* Bit fields: :N with optional < (LSB-first) or > (MSB-first).
* Length-by-reference arrays: [@earlier_field]T.
* TLV records: V(t=u1,l=u2,e=le,h=v|l|a,d=table). Dispatch table
uses tlv.<name>.<hex-tag> in the typedb formats hash.
Architecture highlights:
* Per-instance ReadState (bit_cursor + sibling lookup window).
The bit cursor snap-flushes to the next byte when a non-bit
field follows mid-byte.
* Recursion bound via RzPfCtx::max_depth (default 32).
* Five render modes: text (default), json, cstruct, quiet, dot.
The dot renderer emits a column-aligned record with offset/type/
name/value rows, suitable for direct rendering through Graphviz.
* Positioned diagnostics: RzPfError carries severity, category,
column position, and a human-readable message. Each is also
emitted via RZ_LOG_WARN for backward compatibility.
* Optional RzPfPalette and RzPfRenderOpts threaded through render
for inline ANSI colorisation. NULL palette = canonical text.
* Pointer dereference via RzPfCtx::read_at callback.
The integration test expectations in test/db/cmd/cmd_pf* are
updated in the same commit to match the new text-output format and
the new DOT renderer's column-aligned record layout. Bundling these
keeps the test suite green at this commit.
A transitional rz_type_format_data() shim at the bottom of
pf_parser.c forwards legacy callers (cprint.c, disasm.c) to the new
API; it is removed in the librz/core migration commit.
* librz/arch/types: migrate type DB to new pf format codes
Map every legacy single-character format code in the bundled type
SDB files to its new-DSL equivalent:
b / C -> x1 (1-byte hex)
c -> c (signed char, unchanged)
w -> x2 (2-byte hex LE)
i -> d4 (signed 32-bit decimal LE)
d -> x4 (4-byte hex LE)
q -> x8 (8-byte hex LE)
f -> f4 (32-bit float LE)
F -> F8 (64-bit float BE)
Z -> z(utf16le)
x -> x4
o -> o4
X -> r (hexdump)
n / N -> d4 / u4
Also add new typedef-only types for explicit display formats:
bin{8,16,32,64}_t -> b{1,2,4,8} (binary)
hex{8,16,32,64}_t -> x{1,2,4,8} (hex)
oct{8,16,32,64}_t -> o{1,2,4,8} (octal)
These let users opt into a display representation per field without
changing the underlying scalar size.
Update the corresponding integration test expectations in
test/db/cmd/cmd_avg and test/db/cmd/types: the 'tu', 'tuc', 'tp',
and 'avgp' commands report formats via the SDB, so the migrated
codes appear in their text output (e.g. 'pf 0f4x4 a b' instead of
the legacy 'pf 0fd a b' for a union of float + int).
* librz/core: rewrite pf integration on the new parser API
Remove the legacy rz_type_format_data() bridge and rewrite every
caller to use the new <rz_pf.h> API directly.
librz/core/cprint.c
core_print_format() constructs an RzPfCtx populated from the
RzCore (typedb, big_endian, bits, max_depth) and calls
rz_pf_format() in one shot. The bridge between RzCore I/O and the
pf reader is cprint_pf_read_at(), which forwards to rz_io_nread_at.
The bitmask-to-RzPfMode mapping is cprint_pf_mode().
For DOT mode the format name is passed via RzPfRenderOpts::graph_label
so the dot renderer uses it as the top-level record label.
When scr.color > 0 in TEXT mode, an RzPfPalette is populated from
the active color theme via RzConsPrintablePalette so `pf` output
blends with the rest of Rizin's UI and respects user theme
choices (eco / ec). The palette slots map to theme fields as:
pf field offset <- pal->offset (address column)
pf field name <- pal->fname (symbolic name)
pf endian marker <- pal->meta (metadata tag)
pf hex/number literal <- pal->num (numeric literal)
pf typedb label <- pal->flag (resolved symbol)
reset <- pal->reset
Each slot falls back to its previous hardcoded ANSI escape if the
theme field is unset, preserving prior behaviour for stripped-down
contexts. Other modes (json/cstruct/quiet/dot) ignore the palette.
Closes rizinorg/rizin#782.
librz/core/disasm.c
RZ_META_TYPE_FORMAT case rewritten to resolve via rz_pf_resolve_name
and decode via rz_pf_format() directly. Palette is supplied from
the same theme palette when ds->show_color is true so that
pd @ <struct> blends with the surrounding listing.
librz/include/rz_type.h
Drop the rz_type_format_data() forward declaration. Callers must
use the public <rz_pf.h> surface instead.
librz/core/cmd_descs/cmd_print.yaml
Rewrite the pf command help for the new DSL: sized integers
(case-discriminated endian), special scalars (G GUID, V TLV,
: bits, @ align), strings with encodings, parameterized timestamps,
typed composites (E enum, B bitfield, ? struct), DSL extensions
(length-prefix strings, length-by-reference arrays, inline
bitfields), skip/repeat/pointers, example invocations.
cmd_descs.c is regenerated from the YAML.
* test: pf parser unit and DSL-extension integration tests
Add fresh test coverage for the new pf parser. These are purely
additive: existing integration tests were updated in the parser
and SDB-migration commits so the suite stayed green throughout.
test/unit/test_pf.c (new, 131 tests)
Field-size tables (every fixed-size type) and ctype mapping, every
parse path (sized integers / floats / strings / timestamps /
pointers / structs / enums / bitfields / GUIDs / TLV / alignment /
bits / arrays), every read path (with both fixed and length-by-
reference array counts), every render mode (text / json / cstruct /
quiet / dot), every DSL extension (align, bits MSB/LSB, GUID layout,
length-ref array, length-prefix string, inline bitfield, TLV with
dispatch table), pointer dereference via an in-memory I/O callback,
recursion safety on self-referential structs, lifecycle / NULL
safety, diagnostics (positioned errors, all categories, caret-line
formatter, source capture, verbose parse), ambiguity (each
superficially similar DSL form parses unambiguously), the palette /
colorisation API (issue #782) including a regression test surfaced
by the property-based test harness for input bytes containing ESC,
and DOT-mode coverage (column-aligned layout, single field,
empty/filtered, timestamp values, TLV records).
test/db/cmd/cmd_pf_dsl (new)
8 integration tests targeting the new DSL extensions specifically:
@N alignment, :N bits in MSB and LSB ordering, default-layout GUID,
z[N] length-prefixed string, B4(K=V) inline bitfield, bare V TLV
with defaults, V(t=u2,l=u2,e=be) configured TLV.
test/db/cmd/types_format (new)
Verifies that the new bin{N}_t / hex{N}_t / oct{N}_t typedefs apply
the expected display representation when used inside struct fields
via the 'tp' command, demonstrating the per-base rendering
(binary / decimal / octal / hex).
* doc: pf DSL reference and librz/type README
doc/pf.md (new)
User-facing reference for the pf format DSL covering quick
examples, spec grammar (sized integers with case-discriminated
endian, floats, strings with encodings and length prefixes,
parameterized timestamps, composites -- struct ?, enum E, bitfield B
inline and typed, GUID G, TLV V, raw hexdump r), repetition and
arrays ([N], [@name], {N}, leading 0 for union), padding and
alignment (. skip, @N align, :N bits MSB / LSB), pointer
dereference (*<type>), names grammar (plain vs (typename)name),
output modes (text, json, cstruct, quiet, dot), and diagnostics
(severity / category / position / caret-line formatter; backward-
compatible RZ_LOG_WARN channel).
librz/type/README.md (new)
Subsystem architecture doc for librz/type. Covers public headers,
file map, the pf three-stage pipeline (parse -> read -> render) with
a diagram and entry-point table, per-stage explanations (parser
shape, ReadState scoping rules including bit cursor and sibling
lookup window and snap-flush rule, render mode dispatch including
the DOT column-aligned layout and the graph_label parameter),
typedb integration and recursion bound, error-reporting model,
test coverage summary, and pointer to doc/pf.md as the user-facing
reference.
* librz/type: remove legacy DSL conversion shim
All in-tree callers -- librz/bin/d/, librz/bin/format/, and the
integration tests in test/db/cmd/ -- have been migrated to the new
DSL in the earlier commits of this series. The compatibility shim
that translated bare legacy specifiers (x -> x4, b -> x1, w -> x2,
nN -> uN, NN -> UN, etc.) into the new DSL at the entry of
rz_pf_parse() is no longer load-bearing; this commit deletes it.
- librz/type/pf_parser.c: drop the 267-line
convert_legacy_to_new_dsl() function, the WARN_LEGACY macro,
and the per-parse allocation + free of the converted string.
rz_pf_parse() now walks the caller's spec directly.
- test/db/cmd/cmd_pf,cmd_pf2,cmd_pf_write,cmd_pfd,cmd_pf_new,
metadata: migrate legacy specifiers in CMDS blocks to the new
DSL (b -> x1, w -> x2, q -> x8, i -> d4, bare x -> x4, etc.)
and regenerate EXPECT blocks against the new render where
affected.
The parser now only accepts the new DSL. Any caller still passing
bare legacy specifiers will get "unknown specifier" warnings.
* librz/core,type: add scr.pf.short for delta-offset rendering
When scr.pf.short is enabled, MODE_TEXT rendering shows offsets as
deltas (+<n>) from the format's base address instead of the absolute
hex address. Useful for self-contained struct dumps where the
absolute load address is noise:
$ rizin -e scr.pf.short=true -qc 'wx 00007a452a4b9a02
pf fcb1d4 a b c d' =
0 : a = 4000 [LE]
+4 : b = '*'
+5 : c = 0b0100_1011 [LE]
+6 : d = 666 [LE]
Nested structs reset their delta from the same base, so a child of
`head` at offset +8 still prints as +0 in its own struct block.
The base is derived from the offset of the first top-level value;
callers can override it via RzPfRenderOpts::base_offset.
Includes 5 integration tests in test/db/cmd/cmd_pf_short covering
basic struct, off-mode parity (control), nested struct, concatenated
TLV records, and @-offset usage.
* librz/type,test,doc: add v(N) bitvector type for forensics use
Adds a new specifier `v(N)` to the pf DSL that reads N individual
bits (1..4096) from `ceil(N/8)` bytes and exposes them as N
separate 0/1 scalars. Forensics targets: page-frame bitmaps, NTFS
$Bitmap clusters, ext4 block/inode bitmaps, ELF DT_FLAGS_1, PE
characteristics, ACL bitmasks -- anywhere you want to *see* a
bitmap rather than collapse it to a hex number.
Grammar:
v(N) N-bit bitvector, default MSB-first per byte
v(N,lsb) LSB-first per byte (Intel order)
v(N,msb) explicit MSB-first (DWARF / network order)
Rendering:
text: [ 1 0 1 0 1 0 1 1 | 1 1 0 0 ] (12-bit)
json: {"bit_width":12,"value":"101010111100"}
quiet: 1 0 1 0 1 0 1 1 1 1 0 0
Bitvec fields do not participate in the packed-bit cursor used by
:N; each v(N) reads whole bytes and stands alone. A v(N) field
adjacent to a :N field flushes the partial bit cursor first.
Width is clamped to [1, 4096] with a RANGE diagnostic. The reader
consumes exactly ceil(N/8) bytes, leaving the cursor positioned for
the next field. Verified by the existing tail-field test pattern.
Tests:
test/unit/test_pf.c +10 unit tests (146 total OK)
test/db/cmd/cmd_pf_new +9 cmd tests (440 OK / 9 BR / 0 XX / 4 FX
across all 19 pf suites)
property harness +7 properties (140 unique props, 1000
trials each, 0 failures)
Docs:
doc/pf.md bitvector section with examples
librz/type/README.md cursor-interaction note for v(N) vs :N
* librz/core,type,test: address PR CI feedback for pf rewrite
Five fixes prompted by CI feedback and reviewer comments on the pf
rewrite series:
* librz/core/cmd_descs/cmd_print.yaml: two help-text lines exceeded the
yamllint 120-char ceiling (n1/n2/n4/n8 comment and the deprecation
notes entry). Both now use the same folded-scalar (comment: >) form
already used in cmd_analysis.yaml and cmd_descs.yaml. cmd_descs.c is
regenerated.
* test/db/cmd/cmd_pf_write: drop the BROKEN 'pf xxd print' test. It
relied on '.pf*' (run pf output as rizin commands) which the new
parser no longer emits -- analogous to the '.pfw' removal done
earlier. The companion 'pf xxd print happy' test exercises the same
path through 'pf.' and stays.
* librz/type/pf_parser.c: thread the RzTypeDB through render_cstruct
and render_dot (and their inner helpers) so RZ_PF_ENUM resolution
works in pfc and pfd mode the same way it already does in text mode.
Without this, 'pfc elf_header' emitted '/* 0x00000002 */' and 'pfd'
emitted '|value|0x0000ffff|'; with it the comments and cells carry
the symbolic name (ELFCLASS64, ET_HIPROC, etc).
scalar_text() RZ_PF_BITFIELD: when no inline B4(K=V,...) flags are
present but val->type_name names a typedb enum, walk that enum and
treat each case as a settable bit name. This makes the typedb form
'B (pe_characteristics) flags' decode to
'0x00008140 : IMAGE_DLLCHARACTERISTICS_DYNAMIC_BASE | ...' instead
of just the raw hex -- previously the bitname decoding only worked
on inline B4(...) forms.
* librz/core/cprint.c: cprint_pf_mode() had no mapping for
RZ_PRINT_VALUE, so 'pfv' fell through to TEXT mode and emitted
'0x00000004 = 0x11111111 [LE]' instead of the bare value '0x11111111'
expected by db/esil/arm_16 and existing pfv users. Added an explicit
RZ_PRINT_VALUE -> RZ_PF_MODE_QUIET mapping; QUIET already emits one
bare scalar per field, which is exactly the pfv contract.
* test EXPECTs: cmd_pfd unbreaks its bitfield case (also fixes the
data buffer -- the original used 'wx 0x00008140' which is parsed
byte-by-byte and yields LE u32 0x40810000, not 0x00008140, so no
flags matched). cmd_pf 'PE test' picks up the now-decoded
characteristics and dllCharacteristics bitfields. cmd_pf 'Print
value only' updates to the bare-value pfv contract. test/db/cmd/print
'elf 64bit ls pfc.elf_header' picks up the decoded enum symbols.
Also picks up clang-format-20 reflows in cprint.c, cmd_print.c,
pf_parser.c, and rz_pf.h that the CI's clang-format check flagged.
Also bumps the offset-delta buffer in render_val_text from 16 to 24
bytes so a worst-case ut64 decimal (20 digits + sign + NUL = 22) fits
without tripping -Werror=format-truncation on gcc with -O0.
Also fixes three swapped-argument mu_assert calls in test_pf.c (the
v(N) bitvector tests added in commit 8): the macro signature is
mu_assert(message, test) but those three calls passed (test, message),
which tripped -Werror=format= under -O3 because the int 'test' was
fed to the format string's %s slot.
Adds a second round of regression fixes prompted by a manual audit
against pristine upstream (review feedback after the first pass):
* render_quiet: handle raw byte-sequential types (RZ_PF_UINT128,
RZ_PF_HEXDUMP). scalar_text() has no per-byte case for these so the
default-clause '?' was emitted, making 'pfq Q' and 'pfv Q' print '?'.
render_quiet now mirrors text mode and emits a space-separated hex
stream for raw types.
* Sized pointers (p{2,4,8}): the legacy parser accepted 'p2', 'p4',
'p8' to force a 16/32/64-bit pointer width regardless of ctx.bits;
the new parser only handled bare 'p'. Restored via a new dispatch
branch ahead of bare 'p', recording the override in
RzPfField::bit_width (overloaded -- documented in rz_pf.h). A new
fld_ptr_size() helper consults the override first; read_ptr, the
pointer-deref read path, the STRPTR read path, and
pf_struct_size_impl all use it. 'pf p2p4p8pp2' now consumes
2+4+8+4+2=20 bytes and renders the right value at each position.
* Numeric pointer dereference: '*d4', '*x2', '*u8', etc. (pointer to
fixed-size scalar) only read the pointer itself and never followed
it. The reader now also calls ctx->read_at to fetch the target word
and populate scalars[0]; the text renderer prints both the pointer
literal and the dereferenced value -- '(*0x20) 42' instead of bare
'(*0x20)'. Affects 'pf *d4 ...' and the 'Pointers' / '32 bit twice
then string' tests.
* Pointer-to-struct dereference (*?): pointer-to-struct fields read
the pointer value but stopped there, so 'pf *?' (or typedb forms like
'(troll)Bah' marked with '*') showed only '(*0x30)'. The reader now
recurses through ctx->read_at into a worst-case 4 KiB buffer and
re-invokes read_nested_struct; the renderer prints the nested struct
body after the '(*ptr)' annotation. Recursion is bounded by
ctx->max_depth so cyclic Bah->Bah pointer chains terminate at the
documented limit instead of unbounded recursion. Affects 'nested
struct', 'complex nested struct', and 'flag for nested struct'.
* String pointer (bare 's'): RZ_PF_STRPTR rendering didn't show the
dereferenced target -- output was bare '"hello"' with no pointer
context. Now emits '(*ptr) "string"' (mirroring '*z') when the
deref produced a non-empty body; falls back to bare '""' on
unmapped or empty targets to avoid advertising a phantom pointer.
* render_val_text is_pointer branch: extended to render the nested
struct body for *? pointers and the dereferenced value for numeric
*d4/*x2/*u8 forms.
* Top-level read loop: rz_pf_read passed both 'cur_off + off' (as
read_field's off parameter) AND 'base_addr + cur_off' (as base_addr),
but read_field computes 'val->offset = base_addr + off'. This
double-added cur_off on every iteration past the first, so 'pf 2ic'
reading 5 bytes per iter reported iter 1's nb at offset 0x0a instead
of 0x05, and 'pf 2F' reported iter 1's double at offset 0x10 instead
of 0x08. Data reads were correct; only the displayed/JSON offsets
drifted. The fix: pass base_addr unchanged so val->offset becomes
base_addr + (cur_off + off). Affects 'array obj', 'print n-times a
format', 'pf', 'pf field name', 'JSON output'.
* Bare 'C' (legacy 1-byte unsigned decimal): the legacy parser accepted
'C' as 'print byte as decimal'; the new parser had dropped it. The
'types' test format 'pf fcb1d4C foo bar fool beer plop' was therefore
truncating to 4 fields. Restored as a deprecated alias for 'u1' (with
the standard 'use u1' note).
* pfw dotted-path navigation: the write-mode lookup only matched
top-level field names, so 'pfw gobelin.Buh.first=42' through nested
structs failed with 'field not found'. Replaced with a segment-by-
segment walker that descends through children at each dot. Affects
'write specific element through nested struct'.
* Drop '.pfw' usage from tests. '.pfw' (execute pfw output as rizin
commands) was a legacy convention -- the new pfw writes directly via
rz_io_write_at and emits a human-readable confirmation line, so the
'.' prefix evaluates that confirmation line as a command and fails.
Same direction as the earlier '.pf*' removal. All cmd_pf,
cmd_pf_write, cmd_pf2 tests updated.
* test/db/cmd/cmd_pf 'Register' marked BROKEN with a tracking comment:
the legacy 'r (regname)' looked up CPU registers via
RzPrint::get_register; the new parser repurposes 'r' as raw hex byte
dump and the register-fetch path is gone. Restoring would mean
wiring a new register-lookup hook through RzPfCtx, which is outside
the parser-rewrite scope.
* test/db/cmd/cmd_pf2 'pf F max precision (#13027)' annotated: not a
precision regression. Values differ from the original #13027 expected
output because the legacy 'F' specifier read bytes as LE-then-
reinterpret while the new parser follows the documented 'UPPERCASE
= BE' rule literally; the 17-digit precision -- the actual subject
of #13027 -- is preserved.
EXPECTs updated to match the corrected output across cmd_pf, cmd_pf2,
cmd_pf_write.
* doc/pf.md updated to reflect all DSL and rendering changes:
documents the sized pointers (p/p2/p4/p8), pointer dereference
semantics (numeric, struct, and string), the missing composites
(Q, U, L, n/N), the case-as-endian rule (no standalone endian
directive), the pfw write mode (dotted-path navigation, direct I/O,
no '.pfw'/'.pf*'), and the table of deprecated single-letter
aliases (b, d, o, q, u, i, f, F, w, Z, t, T, X, C).
* librz/type/pf_render.c (new): extracted the render-mode code paths
from pf_parser.c into their own translation unit. Hosts the text /
quiet / JSON / cstruct / DOT renderers, the per-mode helpers
(scalar_text, scalar_json, render_binary, render_guid,
compute_name_width, field_matches, emit_colored*), the RenderCtx
bag, and the public rz_pf_render() + rz_pf_render_json() entry
points. pf_parser.c keeps the parse + read pipeline and shrinks
from 4831 to 3280 lines. Helpers used by both translation units
(is_string_type, is_raw_type, endian_str, pf_vasprintf) move to
pf_parser.h as static inlines.
* librz/type/pf/ (new directory): moved every pf source out of the
librz/type/ top level into a dedicated pf/ subdirectory; only the
legacy format.c stays top-level. meson.build updated with the new
paths and the pf/ include dir.
* Split the monolithic parser into focused sub-grammar TUs, wired
together by the new pf/pf_internal.h (which centralizes the PF_DIAG
diagnostic macro, the shared ReadState, and all cross-TU
declarations):
- pf_parser_string.c string/encoding specs (z / s / Z)
- pf_parser_bitfield.c inline + typed bitfields
- pf_parser_bitvec.c bitvectors v(N)
- pf_parser_array.c array-count resolution
- pf_parser_struct.c nested struct / union reading
pf_parser.c shrinks from 3280 to ~2620 lines and now hosts only the
parse driver, type-spec dispatcher, reader core, context, and the
public utility surface. The TLV TU was migrated off its bespoke
extern/TLV_DIAG block onto the shared header.
* Add an RzStructuredData renderer: pf/pf_render_sd.c with the public
rz_pf_render_sd() (declared in rz_pf.h). (Named _sd, not _sdb: SDB is
Rizin's key/value database, unrelated to RzStructuredData.) It maps the decoded value
vector to the generic key/value document model -- scalars to typed
entries, arrays/bitvectors to arrays, nested structs to sub-maps
with a _type tag, raw/GUID payloads to byte blocks, timestamps to a
formatted string plus a <name>_raw sibling -- so callers get JSON,
YAML, or iterator access for free. Unit and property tests cover it
(see the consistency-pass note below).
* doc/pf.md and librz/type/README.md updated for the new pf/ layout
and the structured-data render mode.
* Consistency pass over librz/type/pf/: deduplicated the structured-data
scalar dispatch (one classifier feeding thin map/array emitters instead
of two parallel ~60-line switches), dropped unused includes
(pf_parser_time.h from the TLV TU, string.h from the struct TU,
rz_endian.h from the SD renderer) and a redundant extern decl now
covered by pf_internal.h, and fixed a dangling-pointer bug in the SD
char path (the classifier returned a pointer into its own by-value
temporary). Expanded coverage: 11 SD unit tests (157 total) and 6 new
theft properties (SD tree non-NULL, JSON structural well-formedness,
top-level object shape, YAML safety, determinism, filter-narrows).
* Split pf_render.c (1548 lines) into per-mode renderer TUs, mirroring
the earlier pf_render_sd.c extraction: pf_render_text.c (text+quiet),
pf_render_json.c (+rz_pf_render_json), pf_render_cstruct.c, and
pf_render_dot.c. A new pf_render.h carries the shared RenderCtx record
plus the cross-TU helper (pf_field_matches, pf_scalar_text,
pf_render_guid) and per-mode entry-point declarations. pf_render.c now
holds only those shared helpers and the rz_pf_render() dispatcher
(~372 lines). meson.build, the file-doc layout listings, and
librz/type/README.md are updated accordingly.
* Second consistency pass: ensured every public RZ_API entry point has a
Doxygen \brief block (added ones for rz_type_format_struct_size and
rz_pf_render_sd) and every file a \file description (added to
pf_parser.h and pf_parser_time.h). Deduplicated the field-size logic:
pf_struct_size_impl's sum and union loops now share one
pf_field_static_size() helper, the enum/bitfield byte-width override is
a single pf_enum_bitfield_width() helper (was inlined 3x), sized-pointer
width routes through the existing fld_ptr_size(), and the inline-bitflag
test shared by the JSON and DOT renderers became the
pf_value_has_inline_bitflags() predicate in pf_render.h. Dropped unused
includes left over from the render split. No behaviour change (157 unit
tests, 146 theft properties, full cmd sweep all green).
* Fix a CodeQL High alert (cpp/integer-multiplication-cast-to-long) in
the scalar-array path selector: the per-element byte delta was computed
as `idx * width` in 32-bit int before being added to the ut64 offset,
so a large array index could overflow int prior to the widening. The
multiplication is now done in ut64 ((ut64)idx * width). No behaviour
change for in-range indices; verified by the path-selection cmd tests.
* Fix a big-endian bug in the structured-data renderer: the bitvector
reader stores each bit in the RzPfScalar union's v_u8 member, but the
SD classifier read it back through v_u64. On little-endian those alias
to the same low byte so it worked by luck; on big-endian (s390x) v_u8
is the high byte of the 64-bit slot, so every set bit read as a large
number and the test_pf_render_sd_bitvec assertion failed. The
classifier now reads v_u8 for RZ_PF_BITVEC, matching the writer; the
other scalar types already read the same width the reader wrote.
* Clean up leftover refactoring-era comments in the renderer TUs: drop
the "Split out of pf_render.c" changelog phrasing from the file-doc
headers (the docs now describe the current layout) and update two
stale references to the pre-rename helper names (field_matches ->
pf_field_matches, render_guid -> pf_render_guid) in comments.
---------
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>