* librz/core,cmd: add pf-aware autocompletion arg types
The `pf` family of commands (`pf`, `pf-`, `pfa`, `pfc`, `pfd`, `pf.`,
`pfn`, `pfo`, `pfs`, `pfv`, `pfw`) takes either a registered named
format, a `<format>.<field>[<idx>]...` path, or a Format Definition
File. None of these were autocompletable: the arg type was always
`RZ_CMD_ARG_TYPE_STRING`, so tab on `pf. <TAB>` did nothing and the
user had to remember every format and field name by hand.
Add three new arg types and wire them to the cmd_descs:
* `RZ_CMD_ARG_TYPE_PF_FORMAT_NAME` enumerates the named formats
from `rz_type_db_format_all()` and filters by the partial prefix.
Used by `pf-`, `pfa`, `pfc`, `pfd`, `pfn`, `pfs`, `pfv`.
* `RZ_CMD_ARG_TYPE_PF_FORMAT_PATH` is the path-aware completer for
`pf.` and `pfw`. Both accept arbitrarily deep paths of the form
`name[.field[<idx>]?]*`, walking through nested struct fields
(e.g. `pf. troll.str[1].two` follows the same path syntax that
`pf_path_navigate` accepts at runtime). The completer:
- Finds the last `.` in the partial input; everything before
it (inclusive) is the committed path, what follows is the
segment being completed.
- Walks each committed `name[N]?` segment in turn, looking up
STRUCT fields' `type_name` in the typedb and re-parsing the
referenced format. Descent past a scalar or an inline struct
(no `type_name`) returns no options.
- Offers the resolved format's field names for the tail.
- Returns nothing when the tail contains `[` or `]` -- the
user is mid-index or mid-descent, where identifier
completion would produce a syntax error if accepted.
- Suppresses the trailing space after each successful
completion so `.` can be typed next without an inserted
space getting in the way of descent.
- Rewrites `res->start` past the last `.` so completion only
replaces the tail; the committed path stays put.
* `RZ_CMD_ARG_TYPE_PF_FDF_FILE` lists the basenames of FDFs found
in the user's home formats dir and the system formats dir,
mirroring the search order used by `pfo` itself; files present
in both locations are reported once via the same `HtSU` de-dup
used in `cmd_print_format_file_handler`. Used by `pfo`.
All three new types are added to `CD_ARG_LAST_TYPES` in
cmd_descs_util.py. The generator sets `RZ_CMD_ARG_FLAG_LAST` on the
final arg of a command whenever that arg's type is in this set;
that flag tells the runtime arg-preprocessor to merge any trailing
whitespace-separated tokens into a single argv slot. This is the
same implicit-FLAG_LAST treatment `RZ_CMD_ARG_TYPE_STRING` already
gets, and it is what keeps invocations like `pfc zd4x8 foo bar cow`
(five tokens, one logical format-with-names argument) working --
without it the cmd parser would reject the extra tokens with "Wrong
number of arguments". The implicit merge does not interfere with
autocompletion: the completer still receives the partial input up
to the cursor and prefix-matches against it, and the format-path
completer's dot/bracket scan is unaffected by whitespace.
The path completer uses a small helper, `pf_path_seg_consume`, that
parses one `name[N]?` segment with safe handling of unterminated `[`,
empty `[]`, non-numeric indices, and end-of-input. `pf_resolve_path_format`
walks all committed segments and returns the RzPfFormat the caller
should complete against; the resolver is the same shape as the
`pf_path_navigate` walker in librz/type/pf/pf_parser.c, except it
operates on RzPfFormat trees (typedb names) rather than RzPfValue
trees (decoded data), so it can run before any read has happened.
The completers live in `cautocmpl.c` next to the existing type-name
completers (`autocmplt_cmd_arg_struct_type` and friends) and follow
the same loop+strncmp pattern. The dispatcher entry for
`RZ_CMD_ARG_TYPE_FOLDER` was missing an explicit `break;` and would
fall through to `default`; harmless, but fixed in passing so the
three new cases sit cleanly above `default:`.
Note on dot-search direction: `rz_sub_str_rchr` is `start..end`
range search returning the FIRST hit, not a right-to-left "find
last" -- the `r` is for "range", not "right". The path completer
needs the last dot, so it walks the buffer backwards itself.
* test/integration: cover pf autocompletion
Thirteen new tests in test_autocmplt.c exercise the three new pf arg
types, including the multi-segment path resolver:
Format name completion:
* `pf_format_name` -- `pfn ut_<TAB>` after registering two formats
confirms both are offered.
Single-segment path completion:
* `pf_format_path` -- three-phase walk through `pf. ut_path<TAB>`,
`pf. ut_path.<TAB>`, `pf. ut_path.cou<TAB>`, covering name-only,
dot-only, and dot-with-prefix. Verifies that `res->end_string`
is empty in the name phase (so `.` can be typed next without an
inserted space) and that `res->start` advances past the dot in
the field phase (so the completion only replaces the field
portion).
* `pfw_format_path` -- the same `<format>.<field>` syntax must
work on the write side too; confirms the PATH completer fires
for `pfw` and is not pf.-specific.
* `pf_format_path_empty` -- bare `pf. <TAB>` lists every
registered format. Snapshots the baseline count first so the
assertion stays robust against any default formats the type DB
might seed.
* `pf_format_path_unknown_name` -- `pf. nonexistent.<TAB>`
returns an empty option list rather than crashing or leaking
diagnostics.
* `pf_format_path_anon_field` -- formats whose fields don't all
have names (e.g. a `.` skip slot) must be iterated safely; the
named fields are offered and the anonymous slot is silently
dropped.
Multi-segment / nested-struct path completion:
* `pf_format_path_nested` -- two-level descent through a STRUCT
field whose `type_name` references another registered format,
parsing the child format and offering its fields.
* `pf_format_path_three_levels` -- three-level descent narrows
correctly: A -> B -> C, then filter C's fields by a prefix.
* `pf_format_path_array_index` -- `pf. troll.str[1].<TAB>` mirrors
the existing cmd_pf2 runtime test; the array index in the
middle segment is parsed and skipped (it doesn't change the
target type).
* `pf_format_path_inside_brackets` -- cursor inside an
unclosed `[` returns no options (mid-index).
* `pf_format_path_after_close_bracket` -- cursor right after `]`
without a trailing `.` also returns no options (mid-descent).
* `pf_format_path_through_scalar` -- descent past a scalar field
is meaningless and returns no options.
The tests use plain `rz_core_new()` (the real cmd_descs already
registers all `pf*` commands), matching the pattern used by
`test_autocmplt_eco_themes`. Format strings use the parser's
"specifier-then-names" form (no internal whitespace in the spec
region) so `rz_pf_parse` produces the expected field count.
* doc,librz/core: align pf docs and `pf?` help with the parser
The standalone reference doc/pf.md and the in-tree `pf?` help (driven
by the details: block in librz/core/cmd_descs/cmd_print.yaml) had
drifted from each other and from what the parser actually accepts.
Both are now consistent with librz/type/pf/pf_parser.c.
Specific corrections:
* `n` family. doc/pf.md claimed `N1`/`N2`/`N4`/`N8` existed as BE
counterparts to `n1`-`n8`, and that bare `n`/`N` defaulted to
`ctx.bits/8`. The parser handles only `n{1,2,4,8}`; all four
forms are context-endian (follow `ctx->big_endian`), and bare
`n` produces "unknown specifier". Rewrite the section to match,
and explain why context-endian is the right choice for header
readers like ELF.
* Deprecation list. The `pf?` "deprecation" note listed `c, s, z`
among the deprecated bare-letter codes -- they are not deprecated
(`c` is the current 1-byte-as-char specifier, `s` is the current
pointer-to-zstring, `z` is the current inline zstring). It was
missing `C, i, Z, X, F, T` which the parser does warn on.
doc/pf.md had `c` in its table marked "unchanged" (so it was
visibly inconsistent with itself) and was missing the `x` row.
Both lists now mirror the parser's PF_DIAG(DEPRECATED) call
sites: b, C, d, f, F, i, o, q, t, T, w, x, X, Z.
* TLV `h=`. `pf?` said "h=v/a (header inclusion)" (two options);
the parser accepts `v` (value only, default), `l` (length covers
len+value), and `a` (length covers tag+len+value). doc/pf.md
already listed all three; help now matches.
* `v(N)` bitvector. doc/pf.md documented this in detail (the
1..4096-bit-wide field type used for things like ELF
`DT_FLAGS_1`, PE characteristics, page-allocation maps), but the
`pf?` help didn't mention it at all. Added to the DSL extensions
section.
* Pointer widths. doc/pf.md documented `p2`/`p4`/`p8` explicitly;
`pf?` only mentioned bare `p` with a "size from ctx.bits"
parenthetical. Help now lists the four forms in one entry.
* GUID layouts. Both docs claimed `G(le)` was "all little-endian",
but the renderer treats `G(le)` and `G(ms)` identically -- D4
(the trailing 8 bytes) is always in buffer order regardless of
layout. Document the actual behaviour rather than the implied
one; this is a description fix, not a code change. Anyone who
wants the byte-reversed-D4 reading can still file it as a
follow-up bug against the renderer in pf_render.c.
cmd_descs.[ch] is regenerated automatically by the custom_target rule
when cmd_print.yaml changes; the .c diff in this commit is the result
of that regeneration (5 lines of comment text inside the existing
detail entries).
format.c can already turn an RzType into a pf format string. Add the
inverse in the same file, next to its counterparts:
rz_type_format_to_c_declaration() parses a pf format string with
rz_pf_parse() and emits an equivalent C struct/union declaration built
from the standard fixed-width types (uint8_t, int32_t, float, ...).
The conversion is structural: it consumes only the parsed RzPfFormat
(field kinds, widths, array counts, pointer flags), never a byte buffer,
so it runs without any target data. Every field kind the engine produces
is mapped; a handful of specifiers with no exact static C form are mapped
best-effort (documented inline): @N alignment is dropped, an
unknown-length inline z string becomes char *, LEB128 widens to its
largest decoded integer, and ?(Name) / E(Name) emit struct/enum
references that must themselves be defined to parse.
Added new 'tdf' command that exposes that conversion to the user
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* librz/type: rewrite pf format parser
Replace the legacy print-format engine in librz/type/format.c with a
clean three-stage pipeline (parse -> read -> render) and a properly
typed DSL.
The new pipeline:
rz_pf_parse(str) source string -> RzPfFormat
rz_pf_read(fmt, buf, len, ctx) RzPfFormat -> RzPfValue[]
rz_pf_render(vals, mode, opts) RzPfValue[] -> printable string
rz_pf_format() one-shot for all of the above
Public API lives in <rz_pf.h> (pulled in transitively via <rz_type.h>).
DSL improvements over the legacy scheme:
* Sized integers with explicit endianness via case:
x4 d4 u4 o4 b4 (LE) vs X4 D4 U4 O4 B4 (BE).
* Sized floats: f2 / f4 / f8 (and BE F2 / F4 / F8).
* Encoding-aware strings: z(utf8) / z(utf16le) / z(utf16be) /
z(utf32le) / z(utf32be) / z(latin1) / z(ebcdic).
* Length-prefixed strings: z[N] reads a N-byte LE prefix then body.
* Parameterized timestamps: t(unix32) / t(unix64) / t(unixms) /
t(unixus) / t(unixns) / t(filetime) / t(dos) / t(hfs) /
t(oletime) / t(webkit) / t(cocoa). ntfs alias maps to filetime.
* GUID: G(ms) (Microsoft mixed-endian, default), G(be), G(le).
* Inline bitfields: B4(FLAG_A=1,FLAG_B=2,FLAG_C=4).
* Alignment: @N advances to next multiple of N.
* Bit fields: :N with optional < (LSB-first) or > (MSB-first).
* Length-by-reference arrays: [@earlier_field]T.
* TLV records: V(t=u1,l=u2,e=le,h=v|l|a,d=table). Dispatch table
uses tlv.<name>.<hex-tag> in the typedb formats hash.
Architecture highlights:
* Per-instance ReadState (bit_cursor + sibling lookup window).
The bit cursor snap-flushes to the next byte when a non-bit
field follows mid-byte.
* Recursion bound via RzPfCtx::max_depth (default 32).
* Five render modes: text (default), json, cstruct, quiet, dot.
The dot renderer emits a column-aligned record with offset/type/
name/value rows, suitable for direct rendering through Graphviz.
* Positioned diagnostics: RzPfError carries severity, category,
column position, and a human-readable message. Each is also
emitted via RZ_LOG_WARN for backward compatibility.
* Optional RzPfPalette and RzPfRenderOpts threaded through render
for inline ANSI colorisation. NULL palette = canonical text.
* Pointer dereference via RzPfCtx::read_at callback.
The integration test expectations in test/db/cmd/cmd_pf* are
updated in the same commit to match the new text-output format and
the new DOT renderer's column-aligned record layout. Bundling these
keeps the test suite green at this commit.
A transitional rz_type_format_data() shim at the bottom of
pf_parser.c forwards legacy callers (cprint.c, disasm.c) to the new
API; it is removed in the librz/core migration commit.
* librz/arch/types: migrate type DB to new pf format codes
Map every legacy single-character format code in the bundled type
SDB files to its new-DSL equivalent:
b / C -> x1 (1-byte hex)
c -> c (signed char, unchanged)
w -> x2 (2-byte hex LE)
i -> d4 (signed 32-bit decimal LE)
d -> x4 (4-byte hex LE)
q -> x8 (8-byte hex LE)
f -> f4 (32-bit float LE)
F -> F8 (64-bit float BE)
Z -> z(utf16le)
x -> x4
o -> o4
X -> r (hexdump)
n / N -> d4 / u4
Also add new typedef-only types for explicit display formats:
bin{8,16,32,64}_t -> b{1,2,4,8} (binary)
hex{8,16,32,64}_t -> x{1,2,4,8} (hex)
oct{8,16,32,64}_t -> o{1,2,4,8} (octal)
These let users opt into a display representation per field without
changing the underlying scalar size.
Update the corresponding integration test expectations in
test/db/cmd/cmd_avg and test/db/cmd/types: the 'tu', 'tuc', 'tp',
and 'avgp' commands report formats via the SDB, so the migrated
codes appear in their text output (e.g. 'pf 0f4x4 a b' instead of
the legacy 'pf 0fd a b' for a union of float + int).
* librz/core: rewrite pf integration on the new parser API
Remove the legacy rz_type_format_data() bridge and rewrite every
caller to use the new <rz_pf.h> API directly.
librz/core/cprint.c
core_print_format() constructs an RzPfCtx populated from the
RzCore (typedb, big_endian, bits, max_depth) and calls
rz_pf_format() in one shot. The bridge between RzCore I/O and the
pf reader is cprint_pf_read_at(), which forwards to rz_io_nread_at.
The bitmask-to-RzPfMode mapping is cprint_pf_mode().
For DOT mode the format name is passed via RzPfRenderOpts::graph_label
so the dot renderer uses it as the top-level record label.
When scr.color > 0 in TEXT mode, an RzPfPalette is populated from
the active color theme via RzConsPrintablePalette so `pf` output
blends with the rest of Rizin's UI and respects user theme
choices (eco / ec). The palette slots map to theme fields as:
pf field offset <- pal->offset (address column)
pf field name <- pal->fname (symbolic name)
pf endian marker <- pal->meta (metadata tag)
pf hex/number literal <- pal->num (numeric literal)
pf typedb label <- pal->flag (resolved symbol)
reset <- pal->reset
Each slot falls back to its previous hardcoded ANSI escape if the
theme field is unset, preserving prior behaviour for stripped-down
contexts. Other modes (json/cstruct/quiet/dot) ignore the palette.
Closes rizinorg/rizin#782.
librz/core/disasm.c
RZ_META_TYPE_FORMAT case rewritten to resolve via rz_pf_resolve_name
and decode via rz_pf_format() directly. Palette is supplied from
the same theme palette when ds->show_color is true so that
pd @ <struct> blends with the surrounding listing.
librz/include/rz_type.h
Drop the rz_type_format_data() forward declaration. Callers must
use the public <rz_pf.h> surface instead.
librz/core/cmd_descs/cmd_print.yaml
Rewrite the pf command help for the new DSL: sized integers
(case-discriminated endian), special scalars (G GUID, V TLV,
: bits, @ align), strings with encodings, parameterized timestamps,
typed composites (E enum, B bitfield, ? struct), DSL extensions
(length-prefix strings, length-by-reference arrays, inline
bitfields), skip/repeat/pointers, example invocations.
cmd_descs.c is regenerated from the YAML.
* test: pf parser unit and DSL-extension integration tests
Add fresh test coverage for the new pf parser. These are purely
additive: existing integration tests were updated in the parser
and SDB-migration commits so the suite stayed green throughout.
test/unit/test_pf.c (new, 131 tests)
Field-size tables (every fixed-size type) and ctype mapping, every
parse path (sized integers / floats / strings / timestamps /
pointers / structs / enums / bitfields / GUIDs / TLV / alignment /
bits / arrays), every read path (with both fixed and length-by-
reference array counts), every render mode (text / json / cstruct /
quiet / dot), every DSL extension (align, bits MSB/LSB, GUID layout,
length-ref array, length-prefix string, inline bitfield, TLV with
dispatch table), pointer dereference via an in-memory I/O callback,
recursion safety on self-referential structs, lifecycle / NULL
safety, diagnostics (positioned errors, all categories, caret-line
formatter, source capture, verbose parse), ambiguity (each
superficially similar DSL form parses unambiguously), the palette /
colorisation API (issue #782) including a regression test surfaced
by the property-based test harness for input bytes containing ESC,
and DOT-mode coverage (column-aligned layout, single field,
empty/filtered, timestamp values, TLV records).
test/db/cmd/cmd_pf_dsl (new)
8 integration tests targeting the new DSL extensions specifically:
@N alignment, :N bits in MSB and LSB ordering, default-layout GUID,
z[N] length-prefixed string, B4(K=V) inline bitfield, bare V TLV
with defaults, V(t=u2,l=u2,e=be) configured TLV.
test/db/cmd/types_format (new)
Verifies that the new bin{N}_t / hex{N}_t / oct{N}_t typedefs apply
the expected display representation when used inside struct fields
via the 'tp' command, demonstrating the per-base rendering
(binary / decimal / octal / hex).
* doc: pf DSL reference and librz/type README
doc/pf.md (new)
User-facing reference for the pf format DSL covering quick
examples, spec grammar (sized integers with case-discriminated
endian, floats, strings with encodings and length prefixes,
parameterized timestamps, composites -- struct ?, enum E, bitfield B
inline and typed, GUID G, TLV V, raw hexdump r), repetition and
arrays ([N], [@name], {N}, leading 0 for union), padding and
alignment (. skip, @N align, :N bits MSB / LSB), pointer
dereference (*<type>), names grammar (plain vs (typename)name),
output modes (text, json, cstruct, quiet, dot), and diagnostics
(severity / category / position / caret-line formatter; backward-
compatible RZ_LOG_WARN channel).
librz/type/README.md (new)
Subsystem architecture doc for librz/type. Covers public headers,
file map, the pf three-stage pipeline (parse -> read -> render) with
a diagram and entry-point table, per-stage explanations (parser
shape, ReadState scoping rules including bit cursor and sibling
lookup window and snap-flush rule, render mode dispatch including
the DOT column-aligned layout and the graph_label parameter),
typedb integration and recursion bound, error-reporting model,
test coverage summary, and pointer to doc/pf.md as the user-facing
reference.
* librz/type: remove legacy DSL conversion shim
All in-tree callers -- librz/bin/d/, librz/bin/format/, and the
integration tests in test/db/cmd/ -- have been migrated to the new
DSL in the earlier commits of this series. The compatibility shim
that translated bare legacy specifiers (x -> x4, b -> x1, w -> x2,
nN -> uN, NN -> UN, etc.) into the new DSL at the entry of
rz_pf_parse() is no longer load-bearing; this commit deletes it.
- librz/type/pf_parser.c: drop the 267-line
convert_legacy_to_new_dsl() function, the WARN_LEGACY macro,
and the per-parse allocation + free of the converted string.
rz_pf_parse() now walks the caller's spec directly.
- test/db/cmd/cmd_pf,cmd_pf2,cmd_pf_write,cmd_pfd,cmd_pf_new,
metadata: migrate legacy specifiers in CMDS blocks to the new
DSL (b -> x1, w -> x2, q -> x8, i -> d4, bare x -> x4, etc.)
and regenerate EXPECT blocks against the new render where
affected.
The parser now only accepts the new DSL. Any caller still passing
bare legacy specifiers will get "unknown specifier" warnings.
* librz/core,type: add scr.pf.short for delta-offset rendering
When scr.pf.short is enabled, MODE_TEXT rendering shows offsets as
deltas (+<n>) from the format's base address instead of the absolute
hex address. Useful for self-contained struct dumps where the
absolute load address is noise:
$ rizin -e scr.pf.short=true -qc 'wx 00007a452a4b9a02
pf fcb1d4 a b c d' =
0 : a = 4000 [LE]
+4 : b = '*'
+5 : c = 0b0100_1011 [LE]
+6 : d = 666 [LE]
Nested structs reset their delta from the same base, so a child of
`head` at offset +8 still prints as +0 in its own struct block.
The base is derived from the offset of the first top-level value;
callers can override it via RzPfRenderOpts::base_offset.
Includes 5 integration tests in test/db/cmd/cmd_pf_short covering
basic struct, off-mode parity (control), nested struct, concatenated
TLV records, and @-offset usage.
* librz/type,test,doc: add v(N) bitvector type for forensics use
Adds a new specifier `v(N)` to the pf DSL that reads N individual
bits (1..4096) from `ceil(N/8)` bytes and exposes them as N
separate 0/1 scalars. Forensics targets: page-frame bitmaps, NTFS
$Bitmap clusters, ext4 block/inode bitmaps, ELF DT_FLAGS_1, PE
characteristics, ACL bitmasks -- anywhere you want to *see* a
bitmap rather than collapse it to a hex number.
Grammar:
v(N) N-bit bitvector, default MSB-first per byte
v(N,lsb) LSB-first per byte (Intel order)
v(N,msb) explicit MSB-first (DWARF / network order)
Rendering:
text: [ 1 0 1 0 1 0 1 1 | 1 1 0 0 ] (12-bit)
json: {"bit_width":12,"value":"101010111100"}
quiet: 1 0 1 0 1 0 1 1 1 1 0 0
Bitvec fields do not participate in the packed-bit cursor used by
:N; each v(N) reads whole bytes and stands alone. A v(N) field
adjacent to a :N field flushes the partial bit cursor first.
Width is clamped to [1, 4096] with a RANGE diagnostic. The reader
consumes exactly ceil(N/8) bytes, leaving the cursor positioned for
the next field. Verified by the existing tail-field test pattern.
Tests:
test/unit/test_pf.c +10 unit tests (146 total OK)
test/db/cmd/cmd_pf_new +9 cmd tests (440 OK / 9 BR / 0 XX / 4 FX
across all 19 pf suites)
property harness +7 properties (140 unique props, 1000
trials each, 0 failures)
Docs:
doc/pf.md bitvector section with examples
librz/type/README.md cursor-interaction note for v(N) vs :N
* librz/core,type,test: address PR CI feedback for pf rewrite
Five fixes prompted by CI feedback and reviewer comments on the pf
rewrite series:
* librz/core/cmd_descs/cmd_print.yaml: two help-text lines exceeded the
yamllint 120-char ceiling (n1/n2/n4/n8 comment and the deprecation
notes entry). Both now use the same folded-scalar (comment: >) form
already used in cmd_analysis.yaml and cmd_descs.yaml. cmd_descs.c is
regenerated.
* test/db/cmd/cmd_pf_write: drop the BROKEN 'pf xxd print' test. It
relied on '.pf*' (run pf output as rizin commands) which the new
parser no longer emits -- analogous to the '.pfw' removal done
earlier. The companion 'pf xxd print happy' test exercises the same
path through 'pf.' and stays.
* librz/type/pf_parser.c: thread the RzTypeDB through render_cstruct
and render_dot (and their inner helpers) so RZ_PF_ENUM resolution
works in pfc and pfd mode the same way it already does in text mode.
Without this, 'pfc elf_header' emitted '/* 0x00000002 */' and 'pfd'
emitted '|value|0x0000ffff|'; with it the comments and cells carry
the symbolic name (ELFCLASS64, ET_HIPROC, etc).
scalar_text() RZ_PF_BITFIELD: when no inline B4(K=V,...) flags are
present but val->type_name names a typedb enum, walk that enum and
treat each case as a settable bit name. This makes the typedb form
'B (pe_characteristics) flags' decode to
'0x00008140 : IMAGE_DLLCHARACTERISTICS_DYNAMIC_BASE | ...' instead
of just the raw hex -- previously the bitname decoding only worked
on inline B4(...) forms.
* librz/core/cprint.c: cprint_pf_mode() had no mapping for
RZ_PRINT_VALUE, so 'pfv' fell through to TEXT mode and emitted
'0x00000004 = 0x11111111 [LE]' instead of the bare value '0x11111111'
expected by db/esil/arm_16 and existing pfv users. Added an explicit
RZ_PRINT_VALUE -> RZ_PF_MODE_QUIET mapping; QUIET already emits one
bare scalar per field, which is exactly the pfv contract.
* test EXPECTs: cmd_pfd unbreaks its bitfield case (also fixes the
data buffer -- the original used 'wx 0x00008140' which is parsed
byte-by-byte and yields LE u32 0x40810000, not 0x00008140, so no
flags matched). cmd_pf 'PE test' picks up the now-decoded
characteristics and dllCharacteristics bitfields. cmd_pf 'Print
value only' updates to the bare-value pfv contract. test/db/cmd/print
'elf 64bit ls pfc.elf_header' picks up the decoded enum symbols.
Also picks up clang-format-20 reflows in cprint.c, cmd_print.c,
pf_parser.c, and rz_pf.h that the CI's clang-format check flagged.
Also bumps the offset-delta buffer in render_val_text from 16 to 24
bytes so a worst-case ut64 decimal (20 digits + sign + NUL = 22) fits
without tripping -Werror=format-truncation on gcc with -O0.
Also fixes three swapped-argument mu_assert calls in test_pf.c (the
v(N) bitvector tests added in commit 8): the macro signature is
mu_assert(message, test) but those three calls passed (test, message),
which tripped -Werror=format= under -O3 because the int 'test' was
fed to the format string's %s slot.
Adds a second round of regression fixes prompted by a manual audit
against pristine upstream (review feedback after the first pass):
* render_quiet: handle raw byte-sequential types (RZ_PF_UINT128,
RZ_PF_HEXDUMP). scalar_text() has no per-byte case for these so the
default-clause '?' was emitted, making 'pfq Q' and 'pfv Q' print '?'.
render_quiet now mirrors text mode and emits a space-separated hex
stream for raw types.
* Sized pointers (p{2,4,8}): the legacy parser accepted 'p2', 'p4',
'p8' to force a 16/32/64-bit pointer width regardless of ctx.bits;
the new parser only handled bare 'p'. Restored via a new dispatch
branch ahead of bare 'p', recording the override in
RzPfField::bit_width (overloaded -- documented in rz_pf.h). A new
fld_ptr_size() helper consults the override first; read_ptr, the
pointer-deref read path, the STRPTR read path, and
pf_struct_size_impl all use it. 'pf p2p4p8pp2' now consumes
2+4+8+4+2=20 bytes and renders the right value at each position.
* Numeric pointer dereference: '*d4', '*x2', '*u8', etc. (pointer to
fixed-size scalar) only read the pointer itself and never followed
it. The reader now also calls ctx->read_at to fetch the target word
and populate scalars[0]; the text renderer prints both the pointer
literal and the dereferenced value -- '(*0x20) 42' instead of bare
'(*0x20)'. Affects 'pf *d4 ...' and the 'Pointers' / '32 bit twice
then string' tests.
* Pointer-to-struct dereference (*?): pointer-to-struct fields read
the pointer value but stopped there, so 'pf *?' (or typedb forms like
'(troll)Bah' marked with '*') showed only '(*0x30)'. The reader now
recurses through ctx->read_at into a worst-case 4 KiB buffer and
re-invokes read_nested_struct; the renderer prints the nested struct
body after the '(*ptr)' annotation. Recursion is bounded by
ctx->max_depth so cyclic Bah->Bah pointer chains terminate at the
documented limit instead of unbounded recursion. Affects 'nested
struct', 'complex nested struct', and 'flag for nested struct'.
* String pointer (bare 's'): RZ_PF_STRPTR rendering didn't show the
dereferenced target -- output was bare '"hello"' with no pointer
context. Now emits '(*ptr) "string"' (mirroring '*z') when the
deref produced a non-empty body; falls back to bare '""' on
unmapped or empty targets to avoid advertising a phantom pointer.
* render_val_text is_pointer branch: extended to render the nested
struct body for *? pointers and the dereferenced value for numeric
*d4/*x2/*u8 forms.
* Top-level read loop: rz_pf_read passed both 'cur_off + off' (as
read_field's off parameter) AND 'base_addr + cur_off' (as base_addr),
but read_field computes 'val->offset = base_addr + off'. This
double-added cur_off on every iteration past the first, so 'pf 2ic'
reading 5 bytes per iter reported iter 1's nb at offset 0x0a instead
of 0x05, and 'pf 2F' reported iter 1's double at offset 0x10 instead
of 0x08. Data reads were correct; only the displayed/JSON offsets
drifted. The fix: pass base_addr unchanged so val->offset becomes
base_addr + (cur_off + off). Affects 'array obj', 'print n-times a
format', 'pf', 'pf field name', 'JSON output'.
* Bare 'C' (legacy 1-byte unsigned decimal): the legacy parser accepted
'C' as 'print byte as decimal'; the new parser had dropped it. The
'types' test format 'pf fcb1d4C foo bar fool beer plop' was therefore
truncating to 4 fields. Restored as a deprecated alias for 'u1' (with
the standard 'use u1' note).
* pfw dotted-path navigation: the write-mode lookup only matched
top-level field names, so 'pfw gobelin.Buh.first=42' through nested
structs failed with 'field not found'. Replaced with a segment-by-
segment walker that descends through children at each dot. Affects
'write specific element through nested struct'.
* Drop '.pfw' usage from tests. '.pfw' (execute pfw output as rizin
commands) was a legacy convention -- the new pfw writes directly via
rz_io_write_at and emits a human-readable confirmation line, so the
'.' prefix evaluates that confirmation line as a command and fails.
Same direction as the earlier '.pf*' removal. All cmd_pf,
cmd_pf_write, cmd_pf2 tests updated.
* test/db/cmd/cmd_pf 'Register' marked BROKEN with a tracking comment:
the legacy 'r (regname)' looked up CPU registers via
RzPrint::get_register; the new parser repurposes 'r' as raw hex byte
dump and the register-fetch path is gone. Restoring would mean
wiring a new register-lookup hook through RzPfCtx, which is outside
the parser-rewrite scope.
* test/db/cmd/cmd_pf2 'pf F max precision (#13027)' annotated: not a
precision regression. Values differ from the original #13027 expected
output because the legacy 'F' specifier read bytes as LE-then-
reinterpret while the new parser follows the documented 'UPPERCASE
= BE' rule literally; the 17-digit precision -- the actual subject
of #13027 -- is preserved.
EXPECTs updated to match the corrected output across cmd_pf, cmd_pf2,
cmd_pf_write.
* doc/pf.md updated to reflect all DSL and rendering changes:
documents the sized pointers (p/p2/p4/p8), pointer dereference
semantics (numeric, struct, and string), the missing composites
(Q, U, L, n/N), the case-as-endian rule (no standalone endian
directive), the pfw write mode (dotted-path navigation, direct I/O,
no '.pfw'/'.pf*'), and the table of deprecated single-letter
aliases (b, d, o, q, u, i, f, F, w, Z, t, T, X, C).
* librz/type/pf_render.c (new): extracted the render-mode code paths
from pf_parser.c into their own translation unit. Hosts the text /
quiet / JSON / cstruct / DOT renderers, the per-mode helpers
(scalar_text, scalar_json, render_binary, render_guid,
compute_name_width, field_matches, emit_colored*), the RenderCtx
bag, and the public rz_pf_render() + rz_pf_render_json() entry
points. pf_parser.c keeps the parse + read pipeline and shrinks
from 4831 to 3280 lines. Helpers used by both translation units
(is_string_type, is_raw_type, endian_str, pf_vasprintf) move to
pf_parser.h as static inlines.
* librz/type/pf/ (new directory): moved every pf source out of the
librz/type/ top level into a dedicated pf/ subdirectory; only the
legacy format.c stays top-level. meson.build updated with the new
paths and the pf/ include dir.
* Split the monolithic parser into focused sub-grammar TUs, wired
together by the new pf/pf_internal.h (which centralizes the PF_DIAG
diagnostic macro, the shared ReadState, and all cross-TU
declarations):
- pf_parser_string.c string/encoding specs (z / s / Z)
- pf_parser_bitfield.c inline + typed bitfields
- pf_parser_bitvec.c bitvectors v(N)
- pf_parser_array.c array-count resolution
- pf_parser_struct.c nested struct / union reading
pf_parser.c shrinks from 3280 to ~2620 lines and now hosts only the
parse driver, type-spec dispatcher, reader core, context, and the
public utility surface. The TLV TU was migrated off its bespoke
extern/TLV_DIAG block onto the shared header.
* Add an RzStructuredData renderer: pf/pf_render_sd.c with the public
rz_pf_render_sd() (declared in rz_pf.h). (Named _sd, not _sdb: SDB is
Rizin's key/value database, unrelated to RzStructuredData.) It maps the decoded value
vector to the generic key/value document model -- scalars to typed
entries, arrays/bitvectors to arrays, nested structs to sub-maps
with a _type tag, raw/GUID payloads to byte blocks, timestamps to a
formatted string plus a <name>_raw sibling -- so callers get JSON,
YAML, or iterator access for free. Unit and property tests cover it
(see the consistency-pass note below).
* doc/pf.md and librz/type/README.md updated for the new pf/ layout
and the structured-data render mode.
* Consistency pass over librz/type/pf/: deduplicated the structured-data
scalar dispatch (one classifier feeding thin map/array emitters instead
of two parallel ~60-line switches), dropped unused includes
(pf_parser_time.h from the TLV TU, string.h from the struct TU,
rz_endian.h from the SD renderer) and a redundant extern decl now
covered by pf_internal.h, and fixed a dangling-pointer bug in the SD
char path (the classifier returned a pointer into its own by-value
temporary). Expanded coverage: 11 SD unit tests (157 total) and 6 new
theft properties (SD tree non-NULL, JSON structural well-formedness,
top-level object shape, YAML safety, determinism, filter-narrows).
* Split pf_render.c (1548 lines) into per-mode renderer TUs, mirroring
the earlier pf_render_sd.c extraction: pf_render_text.c (text+quiet),
pf_render_json.c (+rz_pf_render_json), pf_render_cstruct.c, and
pf_render_dot.c. A new pf_render.h carries the shared RenderCtx record
plus the cross-TU helper (pf_field_matches, pf_scalar_text,
pf_render_guid) and per-mode entry-point declarations. pf_render.c now
holds only those shared helpers and the rz_pf_render() dispatcher
(~372 lines). meson.build, the file-doc layout listings, and
librz/type/README.md are updated accordingly.
* Second consistency pass: ensured every public RZ_API entry point has a
Doxygen \brief block (added ones for rz_type_format_struct_size and
rz_pf_render_sd) and every file a \file description (added to
pf_parser.h and pf_parser_time.h). Deduplicated the field-size logic:
pf_struct_size_impl's sum and union loops now share one
pf_field_static_size() helper, the enum/bitfield byte-width override is
a single pf_enum_bitfield_width() helper (was inlined 3x), sized-pointer
width routes through the existing fld_ptr_size(), and the inline-bitflag
test shared by the JSON and DOT renderers became the
pf_value_has_inline_bitflags() predicate in pf_render.h. Dropped unused
includes left over from the render split. No behaviour change (157 unit
tests, 146 theft properties, full cmd sweep all green).
* Fix a CodeQL High alert (cpp/integer-multiplication-cast-to-long) in
the scalar-array path selector: the per-element byte delta was computed
as `idx * width` in 32-bit int before being added to the ut64 offset,
so a large array index could overflow int prior to the widening. The
multiplication is now done in ut64 ((ut64)idx * width). No behaviour
change for in-range indices; verified by the path-selection cmd tests.
* Fix a big-endian bug in the structured-data renderer: the bitvector
reader stores each bit in the RzPfScalar union's v_u8 member, but the
SD classifier read it back through v_u64. On little-endian those alias
to the same low byte so it worked by luck; on big-endian (s390x) v_u8
is the high byte of the 64-bit slot, so every set bit read as a large
number and the test_pf_render_sd_bitvec assertion failed. The
classifier now reads v_u8 for RZ_PF_BITVEC, matching the writer; the
other scalar types already read the same width the reader wrote.
* Clean up leftover refactoring-era comments in the renderer TUs: drop
the "Split out of pf_render.c" changelog phrasing from the file-doc
headers (the docs now describe the current layout) and update two
stale references to the pre-rename helper names (field_matches ->
pf_field_matches, render_guid -> pf_render_guid) in comments.
---------
Co-authored-by: Anton Kochkov <anton.kochkov@gmail.com>
* Added build files [Capstone to Zydis]
* x86 Analysis [Capstone to Zydis]
* Changed x86 RzIL [Capstone to Zydis]
* Changed asm and arch files [Capstone to Zydis]
* Compilation, build error and test fixes [Capstone to Zydis]
* Test changes [Capstone to Zydis]
* Remaining changes [Capstone to Zydis]
* Added BE support [Capstone to Zydis]
---------
Co-authored-by: tushar3q34 <tushar3q34@gmail.com>
* Fix UB access of NULL context.
During the tests the global RzCons is accessed. Because it was
never initialized all members were 0.
Leading to undefined behavior and breaking the test with ASAN enabled.
* Use correct yaml example for the command described.
* Remove old shell compatibility code.
* Remove dependency tree of rz_core_cmd_subst() and rz_core_cmd_subst_i().
* Remove rz_core_cmd_pipe_old()
* Remove unused rz_core_hack_help()
* Regenerate grammar after removal of legacy_quoted_stmt.
* Bumps Capstone version to newest Capstone next (beyond first v6-Alpha1).
* Fixes leaks
* Fixes build and change to AArch64 and SystemZ compatibility headers.
* Marks M68k test as broken (see commit message).
* Fix AArch64 and SystemZ tests
* Handle op.size == 0 for x86 IL ops
* Add a new test for checking 80-bit floating point operations
* New test `f80_ieee_div_test` tests the division of two 80-bit
floats
* Add SoftFloat 2c as a meson subproject
* Add softfloat code to make the failing test case pass
* Update the hash for the latest softfloat revision
* Implement `rz_float_sqrt` using SoftFloat
* Run the `f80_ieee_div_test` only for x86
* Replace SoftFloat version 2c with 3e
* 3e has less bugs and more features
* Modify the implementation in accordance
* Update SoftFloat revision and add a guard around the 80-bit div test
* Use SoftFloat for add, sub, mul operations as well
* Make rem and mod also use SoftFloat functions
* Also add test for mod and rem, and fix behavior of rem
* Add comment about behavior of mod and rem
* Add comments for tests which have different results for mod and rem
* Simplify usage of loop variable as suggested in review
* Remove unused macro from float.c
* Implement `FMA` and `ROUND` using SoftFloat API
* Add more tests for 80-bit floats
* Change remote to a repository under rizinorg
* Add info about the rounding mode in the Doxygen for rem and mod
* Add comments in tests for rem and mod in `test_float.c`
* Use bitvectors to initialize 80-bit soft floats
* This makes the tests more portable and hence they can be run on
any platform
* Remove guards for f80 tests since they are portable now
* core/cmd: Adjust math commands
* shell: make `?x` commands a parsing failure
We have moved the `?x` commands to `%x`, thus now no command should
start with `?`. This patch makes the parser fail to parse `?x`
strings.
* core/tui: use APIs in panels
* There's no `%q` anymore, use `%=`
RzAnalysisILVM wraps around the low-level RzILVM and enables emulation
of real code from disassembly, rather than raw IL.
Analysis plugins now don't actively initialize the vm anymore, but
return a fully declarative description RzAnalysisILConfig of how to set
up the vm and optionally its initial state.
This also enables multiple IL vms to exist at the same time as plugins
can not mess with the global vm anymore. See the added integration test
for an example.