# LIEF: complete site content > LIEF is a cross-platform C++, Python, and Rust library to parse, analyze, modify, assemble, disassemble, and rewrite ELF, PE, Mach-O, and other executable formats. This file bundles the Markdown version of every page and technical article published on https://lief.re/ into one document. Each document starts with a level-1 heading followed by a metadata list (canonical URL, Markdown URL, authors, dates, tags), and documents are separated by a horizontal rule. Site pages come first (overview, download, about), then articles from newest to oldest. The per-page Markdown files, the discovery map at https://lief.re/llms.txt, the JSON index at https://lief.re/index.json, and the retrieval corpus at https://lief.re/index.chunks.json are generated from the same sources. The technical documentation publishes a separate corpus of guides and API references. Start with its [AI documentation guide](https://lief.re/doc/latest/ai.md) and [discovery map](https://lief.re/doc/latest/llms.txt) to find resolved Markdown pages and documentation indexes. --- # LIEF: Library to Instrument Executable Formats > LIEF is a cross-platform C++, Python, and Rust library to parse, analyze, modify, assemble, disassemble, and rewrite ELF, PE, Mach-O, and other executable formats. - Canonical URL: https://lief.re/ - Markdown: https://lief.re/index.md - Authors: Romain Thomas - Modified: 2026-08-02T11:53:40+02:00 LIEF (Library to Instrument Executable Formats) is an Apache-2.0-licensed, cross-platform library to parse, inspect, modify, and build ELF, PE, Mach-O, and more through one consistent C++, Python, or Rust API. A subset of the API is also exposed in C. ## Key facts - Latest release: LIEF 1.0.0 (2026-07-12) — https://github.com/lief-project/LIEF/releases/tag/1.0.0 - Executable formats: ELF, PE, Mach-O, COFF, and more (see the documentation for the full list); debug formats: DWARF and PDB - LIEF Extended adds an assembler, a disassembler, a DWARF editor, and PDB support on the same object model - Languages: C++ core with Python, Rust, and C APIs - Platforms: Linux, Windows, macOS, Android, and iOS - License: Apache-2.0 - API documentation: https://lief.re/doc/latest/ - Downloads: https://lief.re/download/ - Source code: https://github.com/lief-project/LIEF - Python package: https://pypi.org/project/lief/ - Rust crate: https://crates.io/crates/lief - LIEF Extended: https://extended.lief.re/ ## Documentation for AI assistants Read the AI documentation guide at https://lief.re/doc/latest/ai.md and start with the discovery map at https://lief.re/doc/latest/llms.txt to find guides and API references. In the latest documentation, replace `.html` with `.md` to read resolved Markdown with code examples, every language tab, API signatures, and absolute citation links. - Documentation content index: https://lief.re/doc/latest/index.json - Section and API symbol retrieval: https://lief.re/doc/latest/index.chunks.json - Complete documentation corpus for local indexing: https://lief.re/doc/latest/llms-full.txt (large API references may exceed an assistant's context window). The `latest` channel tracks development; `stable` documents the latest release at https://lief.re/doc/stable/. Match API usage to the installed LIEF version, platform, architecture, and enabled features. ## One library, every layer Move from raw executable file formats to a high-level representation, make precise changes, then write a valid binary back to disk. ### Unified abstraction Work with symbols, sections, relocations, and entry points through shared concepts across executable formats. ### Parse and inspect Open binaries from files or memory and explore their structure without booting a heavyweight reverse-engineering stack. ### Modify and rebuild Add sections, change symbols, patch metadata, and serialize the result with format-aware builders. ## Example: binary insight in a few lines The same concepts stay recognizable whether you work in Python, C++, or Rust. ```python import lief # Parse ELF, PE, or Mach-O with one entry point binary = lief.parse("/usr/bin/ssh") print(binary.format) print(hex(binary.entrypoint)) for section in binary.sections: print(section.name, section.size) # Modify and write a new binary binary.header.entrypoint = 0x401000 binary.write("ssh.patched") ``` ## LIEF Extended Advanced assembly, disassembly, DWARF, and PDB tooling build on the same LIEF model for deeper binary reverse-engineering workflows. - Assembler: assemble x86-64, ARM64, and RISC-V instructions directly from the API. - Disassembler: decode machine code and connect instructions to executable metadata. - DWARF editor: inspect and transform rich debug information with structured APIs. - PDB support: work with Microsoft program databases alongside the binary they describe. Learn more: https://extended.lief.re/ and https://lief.re/doc/latest/extended/intro.html ## Latest articles - [LIEF v1.0.0](https://lief.re/blog/2026-07-13-lief-1-0-0/index.md): 2026-07-13 — LIEF 1.0.0: a new cross-platform Runtime API, faster Rust bindings, stable-ABI and free-threaded Python wheels - [LIEF v0.17.0](https://lief.re/blog/2025-09-14-lief-0-17-0/index.md): 2025-09-14 — LIEF 0.17.0: Binary Ninja and Ghidra plugins, contextual assembly patching, a refactored PE module with TLS, import, and export editing, lief-patchelf, and COFF support. - [LIEF patchelf](https://lief.re/blog/2025-07-13-patchelf/index.md): 2025-07-13 — lief-patchelf: a LIEF-based reimplementation of NixOS patchelf to change the interpreter, add dependencies, and edit RPATH/RUNPATH, with prebuilt binaries. All articles: https://lief.re/blog/ ## Support the project Sponsorship gives LIEF maintainers the time to improve format coverage, respond to ecosystem changes, and build features the community can rely on. - Sponsor LIEF: https://github.com/sponsors/lief-project/ - Discuss a feature: contact@lief.re - Community chat: https://discord.gg/jGQtyAYChJ --- # Download > Download LIEF SDKs and packages for Linux, Windows, macOS, Android, and iOS, or install the Python and Rust bindings. - Canonical URL: https://lief.re/download/ - Markdown: https://lief.re/download/index.md - Authors: Romain Thomas - Modified: 2026-08-02T11:53:40+02:00 Install LIEF from a package manager, download a prebuilt SDK archive, or build it from source. The current release is LIEF 1.0.0 (2026-07-12); release notes: https://github.com/lief-project/LIEF/releases/tag/1.0.0 ## Quick install - Python: `pip install lief` — stable-ABI wheels on PyPI (https://pypi.org/project/lief/), including free-threaded variants where supported. - Rust: `cargo add lief` — the `lief` crate on crates.io (https://crates.io/crates/lief). - From source: clone https://github.com/lief-project/LIEF and follow the compilation guide at https://lief.re/doc/latest/compilation.html. ## Prebuilt packages Official SDK archives and language packages for supported systems and architectures. | Platform | Version | Package | URL | | --- | --- | --- | --- | | Linux | 1.0.0 | SDK x86-64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-x86_64.tar.gz | | Linux | 1.0.0 | SDK x86-64 (Musl) | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-musl-x86_64.tar.gz | | Linux | 1.0.0 | SDK AArch64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-aarch64.tar.gz | | Linux | 1.0.0 | SDK RISC-V 64-bit | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-riscv64.tar.gz | | Linux | 1.0.0 | SDK i686 (Musl) | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-musl-i686.tar.gz | | Linux | 1.0.0 | SDK AArch64 (Musl) | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-musl-aarch64.tar.gz | | Linux | 1.0.0 | SDK RISC-V 64-bit (Musl) | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Linux-musl-riscv64.tar.gz | | Linux | 1.0.0 | Python Packages | https://pypi.org/project/lief/ | | Android | 1.0.0 | SDK ARM | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Android-armv7-a.tar.gz | | Android | 1.0.0 | SDK AArch64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Android-aarch64.tar.gz | | Android | 1.0.0 | SDK x86-64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Android-x86_64.tar.gz | | macOS | 1.0.0 | SDK x86-64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Darwin-x86_64.tar.gz | | macOS | 1.0.0 | SDK AArch64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-Darwin-arm64.tar.gz | | macOS | 1.0.0 | Python Packages | https://pypi.org/project/lief/ | | iOS | 1.0.0 | SDK AArch64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-iOS-arm64.tar.gz | | Windows | 1.0.0 | SDK x86-64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-win64.zip | | Windows | 1.0.0 | SDK x86 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-win32.zip | | Windows | 1.0.0 | SDK ARM64 | https://github.com/lief-project/LIEF/releases/download/1.0.0/LIEF-1.0.0-windows-arm64.zip | | Windows | 1.0.0 | Python packages | https://pypi.org/project/lief/ | | Nightly | main | SDK | https://lief.s3-website.fr-par.scw.cloud/latest/sdk/index.html | | Nightly | main | Python Packages | https://lief.s3-website.fr-par.scw.cloud/latest/lief/index.html | ## Package channels - GitHub releases (SDK archives and release notes): https://github.com/lief-project/LIEF/releases - PyPI (Python wheels): https://pypi.org/project/lief/ - crates.io (Rust crate): https://crates.io/crates/lief - LIEF Extended (assembler, disassembler, DWARF, PDB): https://extended.lief.re/ --- # About > Learn how LIEF became a cross-platform executable-format toolkit, meet its maintainers, and explore the open-source project's milestones. - Canonical URL: https://lief.re/about/ - Markdown: https://lief.re/about/index.md - Authors: Romain Thomas - Modified: 2026-08-02T11:53:40+02:00 LIEF is a cross-platform library that turns executable formats into a structured, scriptable model. It helps security researchers, tool builders, and platform engineers focus on their work instead of format plumbing. ## Project facts - More than ten years of continuous development - More than 1,000,000 installs every month - Four language APIs: C++, Python, Rust, and C - Apache-2.0 license - Maintainer: Romain Thomas ## Project history - 2017 — First public release: LIEF became open source with initial ELF, PE, and Mach-O parsing and modification support. - 2019 — Builders mature: format-aware rewriting expanded with stronger Mach-O modifications and builder APIs. - 2021 — Security metadata: PE Authenticode parsing and verification brought signatures into the unified object model. - 2022 — ELF rewriting at scale: a redesigned ELF builder improved performance, memory behavior, and control over generated binaries. - 2024 — Rust and debug formats: Rust bindings evolved while DWARF and PDB support opened new analysis and tooling workflows. - 2026 — LIEF 1.0: a milestone focused on stability, usability, security, and a new cross-platform Runtime API. ## Community - Contribute on GitHub: https://github.com/lief-project/LIEF - Join the Discord server: https://discord.gg/jGQtyAYChJ - Sponsor LIEF: https://github.com/sponsors/lief-project/ - Contact the maintainers: contact@lief.re - How LIEF works: https://lief.re/doc/latest/intro.html --- # LIEF v1.0.0 > LIEF 1.0.0: a new cross-platform Runtime API, faster Rust bindings, stable-ABI and free-threaded Python wheels - Canonical URL: https://lief.re/blog/2026-07-13-lief-1-0-0/ - Markdown: https://lief.re/blog/2026-07-13-lief-1-0-0/index.md - Authors: Romain Thomas - Published: 2026-07-13T00:00:00Z - Modified: 2026-07-13T00:00:00Z - Tags: release, Runtime API, Rust, Python, LIEF Extended, DWARF, PDB I'm really happy to announce the release of **LIEF 1.0.0**. Compared to previous versions, this release represents an important milestones for this project: stability, usability, and security. LIEF started almost ten years ago. While the journey was driven by adding new features, it also involved correcting a handful of early design decisions that turned out to be wrong and have been corrected, release after release. I won't pretend this version is bug-free or that it does not need further improvements, but it now rests on solid foundations, and the project has been adopted across a range of industries with positive feedback. A special thanks goes to [Quansight](https://quansight.com), which generously sponsored this project. Many thanks as well to [Holepunch](https://holepunch.to/) for their feature-based sponsorship, and to [Calif](https://calif.io/) for reviewing LIEF against frontier models (aka Mythos & GPT‑5‑Cyber). ## Bindings: Rust & Python The LIEF Python wheels are now built against the **stable ABI** and now offer a **free-threaded** variant. Building against the stable ABI lets a single wheel serve multiple interpreter versions, while the free-threaded variant lets you use LIEF with the GIL disabled. This variant is available for Python 3.14 and 3.15 onward. The C++ core is now thread-safe with respect to its few static variables, so it behaves correctly under free-threading. On the Rust side, I refactored the bindings to drop the unmaintained `autocxx` dependency. They are now built directly on top of `cxx`, without any extra wrapper. This change dramatically reduces compilation and iteration time. The crates also moved from `api/rust/cargo/` to `api/rust/crates/`, and the minimum supported Rust version is now `1.85.0`. ## Clang Lifetime Annotations I[^ai-annotations] have started annotating the LIEF API with Clang lifetime annotations `[[clang::lifetimebound]]` to strengthen the compile-time verification of the codebase. These annotations enable the compiler to catch issues like this one: ```cpp std::string lifetime_examples() { Section* text = nullptr; { std::unique_ptr elf = Parser::parse("/bin/ls"); text = elf->get_section(".text"); } return text->name(); } ``` When compiled under strict lifetime-analysis flags, Clang raises the following error: ``` example.cpp:50:12: error: local variable 'elf' does not live long enough [clang-diagnostic-lifetime-safety-use-after-scope] 50 | text = elf->get_section(".text"); | ^~~ example.cpp:51:3: note: local variable 'elf' is destroyed here 51 | } | ^ example.cpp:50:12: note: expression aliases the storage of local variable 'elf' 50 | text = elf->get_section(".text"); | ^~~~~ example.cpp:50:12: note: result of call to 'get_section' aliases the storage of local variable 'elf' 50 | text = elf->get_section(".text"); | ^~~~~~~~~~~~~~~~~~~~~~~~~ example.cpp:52:10: note: later used here 52 | return text->name(); | ^~~~ ``` ## Executable File Format Improvements Mach-O support has been extended to include load commands introduced in recent versions of dyld: `LC_LAZY_LOAD_DYLIB_INFO`, `LC_FUNCTION_VARIANTS`, and `LC_FUNCTION_VARIANT_FIXUPS`. LIEF can now also **create a FAT (universal) binary** from several thin Mach-O binaries, select a specific architecture out of an existing FAT binary, and write **big-endian** Mach-O files. The ELF rewriter has been reworked to reduce the memory footprint of modified binaries: a rewritten ELF is now smaller than in previous releases, and LIEF can safely modify a binary it has *already* modified. Finally, thanks thanks to a security review by Calif, the various parsers are now safer and keep tighter control over how much memory they allocate when facing malformed inputs. ## Supported Platforms This release introduces pre-compiled packages for several new architectures and platforms: - Python wheels are now available for Linux RISC-V 64-bit (including musl-based libc) - SDK and Rust pre-compiled packages are provided for: - Android `x86-64`, `ARM64` & `arm-v7a` - Linux `riscv64gc` (musl), `riscv64a23` (glibc) ## Runtime API One of the most exciting additions in this release is the new **Runtime API**. The motivation came from a recurring need: parsing ELF, Mach-O and PE binaries directly from memory. LIEF already had everything required to parse a binary from a raw pointer thanks to its `BinaryStream` abstraction, but it lacked a friendly, high-level bridge to reach it. The Runtime API is that bridge: ```python import lief for module in lief.runtime.modules(): print(module) binary = module.parse_from_memory() ``` These runtime features go well beyond parsing an executable from memory. They provide cross-platform, cross-language access to inspect and manipulate the process in which LIEF is loaded. To make this concrete, here is a Python example that JIT-compiles and executes a Windows ARM64 "Hello World" using the new Runtime API. First, we allocate a chunk of memory to hold the assembled code: ```python import lief chunk = lief.runtime.Memory.mmap( lief.runtime.Process.page_size, lief.runtime.Memory.ANONYMOUS | lief.runtime.Memory.PRIVATE, lief.runtime.Memory.READ | lief.runtime.Memory.WRITE | lief.runtime.Memory.EXEC, ) ``` Then, we assemble our code straight into the _chunk_ with `assemble(...)`. This function is backed by LIEF's [assembly engine](https://lief.re/doc/latest/extended/assembler/index.html): ```python lief.runtime.assemble( chunk.addr, r""" .text .global win_arm64_hello .align 2 win_arm64_hello: stp x29, x30, [sp, -64]! mov x29, sp stp x19, x20, [sp, 16] stp x21, x22, [sp, 32] // ------------------------------- // Hello World code goes here // ------------------------------- mov x0, xzr ldp x21, x22, [sp, 32] ldp x19, x20, [sp, 16] ldp x29, x30, [sp], 64 ret """, ) ``` While we could implement the "Hello World" as pure, low-level shellcode issuing raw syscalls, we can also leverage the [Contextual Assembly Patching](https://lief.re/doc/latest/extended/assembler/index.html#contextual-assembly-patching) feature to resolve symbols like `GetStdHandle` and `WriteFile` **on the fly**: ```python {linenos=inline hl_lines=[27,28,"5-8"]} class Config(lief.assembly.AssemblerConfig): def resolve_symbol(self, name: str) -> int | None: kernel32 = lief.runtime.windows.dlopen("kernel32.dll") match name: case "GetStdHandle": return lief.to_int(kernel32.dlsym("GetStdHandle")) case "WriteFile": return lief.to_int(kernel32.dlsym("WriteFile")) case _: return None config = Config() lief.runtime.assemble( chunk.addr, r""" .text .global win_arm64_hello .align 2 win_arm64_hello: stp x29, x30, [sp, -64]! mov x29, sp stp x19, x20, [sp, 16] stp x21, x22, [sp, 32] // ------------------------------- ldr x19, =GetStdHandle ldr x20, =WriteFile // ------------------------------- mov x0, xzr ldp x21, x22, [sp, 32] ldp x19, x20, [sp, 16] ldp x29, x30, [sp], 64 ret """, config, ) ``` The same mechanism can also resolve **data** symbols. Let's expose our message buffer and its length through the configuration, then actually call the two functions to print the string: ```python {linenos=inline hl_lines=[1,2,"12-15",36,37,"39-48"]} msg = b"Hello World\n" ctype_msg = ctypes.create_string_buffer(msg) class Config(lief.assembly.AssemblerConfig): def resolve_symbol(self, name: str) -> int | None: kernel32 = lief.runtime.windows.dlopen("kernel32.dll") match name: case "GetStdHandle": return lief.to_int(kernel32.dlsym("GetStdHandle")) case "WriteFile": return lief.to_int(kernel32.dlsym("WriteFile")) case "var_msg": return ctypes.addressof(ctype_msg) case "var_msg_len": return len(msg) case _: return None config = Config() lief.runtime.assemble( chunk.addr, r""" .text .global win_arm64_hello .align 2 win_arm64_hello: stp x29, x30, [sp, -64]! mov x29, sp stp x19, x20, [sp, 16] stp x21, x22, [sp, 32] // ------------------------------- ldr x19, =GetStdHandle ldr x20, =WriteFile ldr x21, =var_msg ldr w22, =var_msg_len // GetStdHandle(STD_OUTPUT_HANDLE=-11) -> x0 mov w0, -11 blr x19 // WriteFile(x0, msg, len, &written, NULL) mov x1, x21 mov w2, w22 add x3, sp, 48 mov x4, xzr blr x20 // ------------------------------- mov x0, xzr ldp x21, x22, [sp, 32] ldp x19, x20, [sp, 16] ldp x29, x30, [sp], 64 ret """, config, ) ``` Once the code is assembled in memory, we flush the instruction cache, flip the region to read-execute, and call it like a regular function: ```python lief.runtime.assemble(chunk.addr, "...", config) chunk.cache_flush() chunk.make_rx() void_void_func = ctypes.CFUNCTYPE(None) # void(*)() hello_jit = void_void_func(chunk.addr) hello_jit() # prints "Hello World" ``` The same example is available for the Rust bindings and the C++ API, and you'll find the Linux, Android, and macOS equivalents in the [runtime documentation](https://lief.re/doc/latest/runtime/intro.html). The Runtime API also works the other way around: you can disassemble live code directly from memory. Here is a Rust example that disassembles a function and flags every instruction that touches the RISC-V stack pointer (`x2`): ```rust use lief::assembly::{Instructions, riscv::{Operands, Reg}}; fn say_hello() { println!("Hello World"); } fn main() { for inst in lief::runtime::disassemble((say_hello as usize).try_into().unwrap()) { println!("{inst}"); let Instructions::RiscV(riscv) = &inst else { continue; }; let uses_sp = riscv .operands() .any(|op| matches!(op, Operands::Mem(mem) if mem.base() == Reg::X2)); if uses_sp { println!("{inst} is a memory operation that uses x2 (sp)"); } } } ``` **Runtime API** You can explore the rest of the Runtime API in the [runtime documentation](https://lief.re/doc/latest/runtime/intro.html). Because these features go beyond executable formats, they are only activated in the [extended](https://lief.re/doc/latest/extended/intro.html) version of LIEF. However, you can also compile LIEF from source with the Runtime API enabled. ## LIEF Extended **C++ Standard** [LIEF Extended](https://extended.lief.re) and its components are now built using **C++23**. The supported platforms are: - Linux: Ubuntu 24.04, Fedora 40, RHEL 10 / CentOS 10, Debian 13: `ARM64, x86-64` (`RISC-V` soon) - macOS 15.0+ (Sequoia): `ARM64 & x86-64` - Windows 11+: `ARM64 & x86-64` - Android API 30+: `ARM64 & x86-64` (including Python wheels) This change **does not** affect the LIEF core library, which remains compatible with C++11 and only requires a C++17 compiler to build. ### DWARF & PDB `->` C/C++ LIEF Extended can now generate C/C++ declarations from DWARF and PDB debug information. For example, consider this [libdexprotector.so](https://www.romainthomas.fr/post/26-01-dexprotector/) binary, enriched with DWARF debug info recovered via reverse engineering. In Binary Ninja, it looks like this: ![Technical diagram](https://lief.re/blog/2026-07-13-lief-1-0-0/bn-1.svg) Calling `to_decl()` on a DWARF function, variable, or type generates its C/C++ declaration: ```python import lief elf = lief.ELF.parse("libdexprotector.so.dwarf") dwarf = elf.debug_info linker64_r_debug = dwarf.find_variable("linker64_r_debug") print(linker64_r_debug.to_decl()) ``` Which outputs: ```cpp /* * pointer to the r_debug structure defined in the linker(64) * Addr: 0xabc8 * size: 0x0008 */ static struct r_debug_t *linker64_r_debug; ``` **C/C++ Declaration** Note that the Binary Ninja comment, the address, and the `sizeof` of the variable are also generated. You can control the output format using the configuration options accepted by `to_decl`. `to_decl` also works with types: ```python {linenos=inline hl_lines=["9-10"]} import lief elf = lief.ELF.parse("libdexprotector.so.dwarf") dwarf = elf.debug_info linker64_r_debug = dwarf.find_variable("linker64_r_debug") print(linker64_r_debug.to_decl()) ptr_type: lief.dwarf.types.Pointer = linker64_r_debug.type print(ptr_type.underlying_type.to_decl()) ``` Which outputs: ```cpp struct r_debug_t { int r_version; char __padding1__[4]; struct link_map *r_map; Elf64_Addr r_brk; enum { RT_CONSISTENT = 0U, RT_ADD = 1U, RT_DELETE = 2U } r_state; char __padding4__[4]; Elf64_Addr r_ldbase; } ``` You can also configure the output to include field offsets: ```diff {linenos=inline style=pastie} import lief elf = lief.ELF.parse("libdexprotector.so.dwarf") dwarf = elf.debug_info linker64_r_debug = dwarf.find_variable("linker64_r_debug") print(linker64_r_debug.to_decl()) ptr_type: lief.dwarf.types.Pointer = linker64_r_debug.type - print(ptr_type.underlying_type.to_decl()) + opt = lief.DeclOpt() + opt.show_field_offsets = True + print(ptr_type.underlying_type.to_decl(opt)) ``` ```cpp struct r_debug_t { /* 0x00 */ int r_version; /* 0x04 */ char __padding1__[4]; /* 0x08 */ struct link_map *r_map; /* 0x10 */ Elf64_Addr r_brk; /* 0x18 */ enum { RT_CONSISTENT = 0U, RT_ADD = 1U, RT_DELETE = 2U } r_state; /* 0x1c */ char __padding4__[4]; /* 0x20 */ Elf64_Addr r_ldbase; } ``` A function containing basic block comments like this: [![BinaryNinja - dp_derive_key](https://lief.re/blog/2026-07-13-lief-1-0-0/bn_screenshot.webp)](https://lief.re/blog/2026-07-13-lief-1-0-0/bn_screenshot.webp) is translated as follows: ```cpp /* * Address: 0x063c */ r_debug_t *dp_derive_key(key_t *key) { /* * Stack addr: -0x00d0 * size: 0x0080 */ struct_4 var_d0; /* Start: 0x000988 */ { /* Start: 0x00098c */ { //This block derives the key based on the assembly of r_debug.r_brk. // //Frida hooks this function such as the regular "ret" is transformed by a trampoline } /* End: 0x000990 */ } /* End: 0x0009a4 */ } ``` The same feature works with PDB files: ```python pdb = lief.pdb.load("ntdll.pdb/6192BFDB9F04442995FFCB0BE95172E14/ntdll.pdb") peb = pdb.find_type("_PEB") opt = lief.DeclOpt() opt.show_field_offsets = True print(peb.to_decl(opt)) ``` ```cpp struct _PEB { /* 0x00 */ unsigned char InheritedAddressSpace; /* 0x01 */ unsigned char ReadImageFileExecOptions; /* 0x02 */ unsigned char BeingDebugged; /* 0x03 */ unsigned char BitField; /* 0x03 */ unsigned char ImageUsesLargePages : 1; /* 0x03 */ unsigned char IsProtectedProcess : 1; /* 0x03 */ unsigned char IsLegacyProcess : 1; /* 0x03 */ unsigned char IsImageDynamicallyRelocated : 1; /* 0x03 */ unsigned char SkipPatchingUser32Forwarders : 1; /* 0x03 */ unsigned char SpareBits : 3; /* 0x08 */ void Mutant; /* 0x10 */ void ImageBaseAddress; /* 0x18 */ struct _PEB_LDR_DATA *Ldr; /* 0x20 */ struct _RTL_USER_PROCESS_PARAMETERS *ProcessParameters; /* 0x28 */ void SubSystemData; /* 0x30 */ void ProcessHeap; [...] /* 0x378 */ unsigned long TracingFlags; /* 0x378 */ unsigned long HeapTracingEnabled : 1; /* 0x378 */ unsigned long CritSecTracingEnabled : 1; /* 0x378 */ unsigned long SpareTracingBits : 30; } ``` ### RISC-V, MIPS, eBPF & PowerPC The disassembler can now expose the high-level semantics of instruction operands for `RISC-V`, `MIPS`, `eBPF` and `PowerPC`. Each operand is surfaced as a `Register`, `Immediate`, `Memory` or `PCRelative` object: ```python import lief elf = lief.parse("hello.bpf.o") section = elf.get_section("tp/syscalls/sys_enter_write") instructions = list(elf.disassemble_from_bytes(bytes(section.content))) for inst in instructions: print(inst) for idx, op in enumerate(inst.operands): match op: case lief.assembly.ebpf.operands.Register(value=reg): print(f" op[{idx}]: Register -> {reg}") case lief.assembly.ebpf.operands.Immediate(value=imm): print(f" op[{idx}]: Immediate -> {imm}") case lief.assembly.ebpf.operands.Memory(): print(f" op[{idx}]: Memory -> {op}") case lief.assembly.ebpf.operands.PCRelative(value=target): print(f" op[{idx}]: PCRelative -> {target}") ``` ## Last Word The first internal version of LIEF was released in July 2016. Ten years later, the project is still alive and keeps growing, supported by a large and active community of users and contributors. Thank you to everyone who has been a part of this journey. The detailed changelog is available here: [Changelog](https://lief.re/doc/latest/changelog.html) [ ![Quansight](https://lief.re/blog/2026-07-13-lief-1-0-0/quansight.webp) ](https://quansight.com/) [^ai-annotations]: The initial pass over the codebase to add these annotations was bootstrapped with the help of AI, then reviewed and adjusted by hand. --- # LIEF v0.17.0 > LIEF 0.17.0: Binary Ninja and Ghidra plugins, contextual assembly patching, a refactored PE module with TLS, import, and export editing, lief-patchelf, and COFF support. - Canonical URL: https://lief.re/blog/2025-09-14-lief-0-17-0/ - Markdown: https://lief.re/blog/2025-09-14-lief-0-17-0/index.md - Authors: Romain Thomas - Published: 2025-09-14T00:00:00Z - Modified: 2025-09-14T00:00:00Z - Tags: release, PE, COFF, assembler, Binary Ninja, Ghidra, patchelf This new version of LIEF introduces several improvements and features that expand the scope of LIEF's use cases. ## Reverse Engineering Plugins Reverse engineering frameworks like Binary Ninja and Ghidra provide excellent support for analyzing instructions and functions. However, they might lack an in-depth analysis of all structures associated with executable formats. For example, the latest version of Ghidra (`11.4.2`) and Binary Ninja (`5.1.8104`) are not able to accurately process Windows ARM64EC binaries, which combine ARM64 code with x86_64. ARM64EC binaries use specific structures such as `IMAGE_ARM64EC_METADATA`, that are not (yet) recognized by most of the reverse engineering frameworks: ![Comparison: before](https://lief.re/blog/2025-09-14-lief-0-17-0/chpe_metadata_before.svg) ![Comparison: after](https://lief.re/blog/2025-09-14-lief-0-17-0/chpe_metadata_after.svg) This `IMAGE_ARM64EC_METADATA` structure contains the `ExtraRFE` attribute, which is used to reference an exception table that is specific to the ARM64EC: ![Comparison: before](https://lief.re/blog/2025-09-14-lief-0-17-0/rdata_before.svg) ![Comparison: after](https://lief.re/blog/2025-09-14-lief-0-17-0/rdata_after.svg) As for the `x86_64` architecture, this table can be used to increase the coverage of functions recognized by BinaryNinja: ![Comparison: before](https://lief.re/blog/2025-09-14-lief-0-17-0/featmap_before.svg) ![Comparison: after](https://lief.re/blog/2025-09-14-lief-0-17-0/featmap_after.svg) The source code of these plugins is located in the main LIEF repository under the following directories: - [plugins/binaryninja](https://github.com/lief-project/LIEF/tree/main/plugins/binaryninja) - [plugins/ghidra](https://github.com/lief-project/LIEF/tree/main/plugins/ghidra) This new release of LIEF introduces official and maintained support for both Binary Ninja and Ghidra, enhancing the analysis capabilities and type definitions for these frameworks. ## Contextual Assembly Patching **Note** The feature is only available in [LIEF Extended](https://lief.re/doc/latest/extended/intro.html) When patching assembly code, we frequently need to refer to external data or functions within our assembly listing. For instance, we might want to naturally write: ```python elf = lief.ELF.parse("libdexprotector.so") elf.assemble( elf.get_function_address("libdp_init"), """ adrp x4, g_protections_conf add x4, x4, :lo12:g_protections_conf strb wzr, [x4, 0xe] // Disable debug check. """ ) ``` In this code snippet, we assume that the address of `g_protections_conf` is known. With the latest release, the assembler engine has introduced support for **dynamically** resolving symbols referenced in an assembly listing. This functionality works by providing an additional configuration parameter: ```python {linenos=inline hl_lines=["1-11", 22]} class Config(lief.assembly.AssemblerConfig): def __init__(self, target: lief.Binary): super().__init__() self._target = target def resolve_symbol(self, name: str) -> int | None: dwarf_info: lief.dwarf.DebugInfo = self._target.debug_info if var := dwarf_info.find_variable(name): print(f"'{name}' is located at address: 0x{var.address:016x}") return var.address return super().resolve_symbol(name) elf = lief.ELF.parse("libdexprotector.so") config = Config(elf) elf.assemble( elf.get_function_address("libdp_init"), """ adrp x4, g_protections_conf add x4, x4, :lo12:g_protections_conf strb wzr, [x4, 0xe] // Disable debug check. """, config ) ``` The logic of the `resolve_symbol` function depends on DWARF information that is generated by BinaryNinja (see: [ BinaryNinja - DWARF Plugin](https://lief.re/doc/latest/plugins/binaryninja/dwarf/index.html#export-as-dwarf)) After patching, we can see the final results, which show that the `adrp` instructions have been correctly generated to access the `g_protections_conf.debug_check` variable: ![Comparison: before](https://lief.re/blog/2025-09-14-lief-0-17-0/dxp_lhs.svg) ![Comparison: after](https://lief.re/blog/2025-09-14-lief-0-17-0/dxp_rhs.svg) For more details about the API, you can check: https://lief.re/doc/latest/extended/assembler/index.html#contextual-assembly-patching ## PE Refactoring ![MSVC PE import table layout showing the IAT, ILT, headers, and string table](https://lief.re/blog/2025-09-14-lief-0-17-0/msvc_layout.webp) LIEF's PE module has been significantly refactored, including improvements to the parser, builder, and documentation. These improvements bring this format to a level of maturity comparable to other formats (ELF/Mach-O). The most notable enhancements involve the ability to modify TLS, as well as manage imports and exports. It also provides support for ARM64EC and ARM64X binaries. For more details, please refer to: - this blog post: https://lief.re/blog/2025-02-16-arm64ec-pe-support/ - this dedicated changelog: https://lief.re/doc/latest/changelog/pe-0-17-0.html#pe-0170-changelog ## lief-patchelf For this release, I started to bootstrap LIEF-based tools (mostly CLI) that aim to provide specific functionalities using LIEF. The first tool of this bootstrap is `lief-patchelf`, which provides a drop-in replacement for the well-known [NixOS/patchelf](https://github.com/NixOS/patchelf). You can find insight about this tool in this blog post: https://lief.re/blog/2025-07-13-patchelf/ ## COFF Format The COFF format is now supported by LIEF, but it does not support modifications yet. The API is really similar to the other formats and available in C++/Rust/Python: ```python import lief coff: lief.COFF.Binary = lief.COFF.parse(r"C:\Users\romain\test.obj") # Access symbols and aux info for symbol in coff.symbols: print(symbol.name) for aux in symbol.auxiliary_symbols: assert str(aux) == """ AuxiliaryCLRToken { Aux Type: 1 Reserved: 1 Symbol index: 10 Symbol: ??0CppInlineNamespaceAttribute@?A0xb81de522@vc.cppcli.attributes@@$$FQE$AAM@PE$[...] Rgb reserved: +---------------------------------------------------------------------+ | 00 00 00 00 00 00 00 00 00 00 00 00 | ............ | +---------------------------------------------------------------------+ } """ # Disassembler support for inst in coff.disassemble("?foo@@YAHHH@Z") print(inst) ``` The documentation for this format is here: https://lief.re/doc/latest/formats/coff/index.html ## Final Word Version `0.17.0` is the **latest** release in the `0.X.Y` series. With significant improvements to the PE format and the global scale at which LIEF is used, I believe the project is now ready for its upcoming `1.0` version. The complete changelog is here: https://lief.re/doc/latest/changelog.html Happy LIEF, Romain [^llvm-engine]: Based on LLVM 21: https://lief.re/doc/latest/extended/intro.html#lief-extended-llvm --- # LIEF patchelf > lief-patchelf: a LIEF-based reimplementation of NixOS patchelf to change the interpreter, add dependencies, and edit RPATH/RUNPATH, with prebuilt binaries. - Canonical URL: https://lief.re/blog/2025-07-13-patchelf/ - Markdown: https://lief.re/blog/2025-07-13-patchelf/index.md - Authors: Romain Thomas - Published: 2025-07-13T00:00:00Z - Modified: 2025-07-13T00:00:00Z - Tags: ELF, patchelf, tooling, Linux ![Featuring Image](https://lief.re/blog/2025-07-13-patchelf/featured.webp) ## Introduction [patchelf](https://github.com/NixOS/patchelf) is a utility developed by the NixOS community to facilitate modifications of some parts of an ELF binary. The features offered by patchelf include, but are not limited to: ```bash # Change the interpreter path $ patchelf --set-interpreter /lib/my-ld-linux.so.2 my-program # Add a new dependency $ patchelf --add-needed libfoo.so.1 my-program # Change RUN/RPATH values $ patchelf --set-rpath /opt/my-libs/lib:/other-libs my-program ``` In addition to its use in NixOS, many projects leverage patchelf to perform post-compilation modifications. For example, [pypa/auditwheel](https://github.com/pypa/auditwheel) uses `patchelf` to change the rpath value and/or change the names of `DT_NEEDED` libraries. ## LIEF & Patchelf Unlike LIEF, patchelf is not a library, and the current version (0.18.0) can only be accessed via the command line. The source code of patchelf is relatively small (around 3,000 lines of code) and is focused on specific modifications such as renaming/removing `DT_NEEDED` entries and changing/removing/adding `DT_RPATH/DT_RUNPATH`. When it comes to design decisions for modifying ELF binaries, there is a significant difference between patchelf and LIEF: **Section vs Segments** ** Patchelf relies on the section layout to modify and rewrite ELF binaries, while LIEF uses segments representation. This distinction makes patchelf less reliable for ELF binaries that have a non-standard section layout, whereas the representation used by LIEF corresponds to how binaries are loaded, making it less prone to errors during modifications. ** ## LIEF-based patchelf LIEF already had (almost) all the internal features to provide the same functionalities as `patchelf`. However, LIEF does not offer a comprehensive command-line tool like patchelf does. To create a drop-in replacement for patchelf using the LIEF backend, we had to ensure compatibility with the current state of patchelf, including: 1. A command-line interface 2. Completion for `zsh` (at least) 3. A man-page Thanks to the recent LIEF's Rust bindings, we can leverage the Rust ecosystem to simplify the development of this command-line tool: - The crate [clap-rs/clap](https://docs.rs/clap/latest/clap/) provides all the features to create a command-line interface with a man-page and completion for the main shells. - Cargo and its dependencies management are convenient for generating an executable. ![Rust, clap-rs, and LIEF components used to build lief-patchelf](https://lief.re/blog/2025-07-13-patchelf/bindings.webp) The source code is available here: [lief-project/LIEF/tools/lief-patchelf](https://github.com/lief-project/LIEF/tree/main/tools/lief-patchelf) and it is worth mentioning that this LIEF-based implementation has the **exact** same command line interface as the original `patchelf`: ```bash $ lief-patchelf --help Patchelf based on LIEF Usage: lief-patchelf [OPTIONS] [filenames]... Arguments: [filenames]... Options: --set-interpreter Change the dynamic loader ('ELF interpreter') of executable given to INTERPRETER. --page-size Uses the given page size instead of the default --print-interpreter Prints the ELF interpreter of the executable. (e.g. `/lib64/ld-linux-x86-64.so.2`) --print-os-abi Prints the OS ABI of the executable (`EI_OSABI` field of an ELF file). ``` Moreover, it passes[^test-note] all the original patchelf's test suite except one test related to the `IA-64` architecture support: ``` PASS: set-interpreter-same.sh PASS: no-rpath-alpha.sh PASS: no-rpath-amd64.sh PASS: no-rpath-armel.sh PASS: no-rpath-armhf.sh PASS: no-rpath-hurd-i386.sh PASS: no-rpath-i386.sh FAIL: no-rpath-ia64.sh PASS: no-rpath-kfreebsd-amd64.sh PASS: no-rpath-kfreebsd-i386.sh PASS: no-rpath-mips.sh PASS: no-rpath-mipsel.sh PASS: no-rpath-powerpc.sh PASS: no-rpath-s390.sh PASS: no-rpath-sh4.sh PASS: no-rpath-sparc.sh ============================================================================ Testsuite summary for patchelf 0.18.0 (based on LIEF 0.17.0) ============================================================================ # TOTAL: 58 # PASS: 55 # SKIP: 2 # XFAIL: 0 # FAIL: 1 # XPASS: 0 # ERROR: 0 ``` ## Download If you want to test this new LIEF-based implementation, you can download the compiled version here: * [`lief-tools-aarch64-apple-darwin.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-aarch64-apple-darwin.zip) * [`lief-tools-aarch64-pc-windows-msvc.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-aarch64-pc-windows-msvc.zip) * [`lief-tools-aarch64-unknown-linux-gnu.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-aarch64-unknown-linux-gnu.zip) * [`lief-tools-aarch64-unknown-linux-musl.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-aarch64-unknown-linux-musl.zip) * [`lief-tools-i686-unknown-linux-musl.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-i686-unknown-linux-musl.zip) * [`lief-tools-x86_64-apple-darwin.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-x86_64-apple-darwin.zip) * [`lief-tools-x86_64-pc-windows-msvc.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-x86_64-pc-windows-msvc.zip) * [`lief-tools-x86_64-unknown-linux-gnu.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-x86_64-unknown-linux-gnu.zip) * [`lief-tools-x86_64-unknown-linux-musl.zip`](https://lief.re/blog/2025-07-13-patchelf/lief-tools-x86_64-unknown-linux-musl.zip) [^test-note]: Since LIEF performs in-depth modifications and can raise different error messages, I had to make minor changes on some checks: [patchelf-test-lief.diff](https://gist.github.com/romainthomas/2332cddd164cf8bfa3fc65cd61496c9a) --- # DWARF as a Shared Reverse Engineering Format > Create DWARF debug information with the LIEF Extended DWARF editor and share reverse-engineered types and functions between Binary Ninja and Ghidra plugins. - Canonical URL: https://lief.re/blog/2025-05-27-dwarf-editor/ - Markdown: https://lief.re/blog/2025-05-27-dwarf-editor/index.md - Authors: Romain Thomas - Published: 2025-05-27T00:00:00Z - Modified: 2025-05-27T00:00:00Z - Tags: DWARF, LIEF Extended, reverse-engineering, Binary Ninja, Ghidra ![Featuring Image](https://lief.re/blog/2025-05-27-dwarf-editor/featured.webp) ## Introduction When reverse engineering binaries, we could want, at some point, to share the reverse-engineered information with others. The DWARF format, originally designed to hold debug information associated with the original source code, is also well-suited for storing reverse-engineered information such as structures and function names. This blog post introduces a new API in *LIEF extended* to create DWARF files. It also introduces two plugins for Ghidra and BinaryNinja to export binary analysis into DWARF. ## Creating DWARF with LIEF (extended) [LIEF extended](https://lief.re/doc/latest/extended/intro.html) now provides a comprehensive API to create DWARF files. This API is available in Python, Rust, and C++ and it looks like this: ```python import lief elf = lief.ELF.parse("./libd5A7BCF0524B8.so") editor: lief.dwarf.Editor = lief.dwarf.Editor.from_binary(elf) unit: lief.dwarf.editor.CompilationUnit = editor.create_compilation_unit() unit.set_producer("Generated by LIEF (LLVM backend)") func: lief.dwarf.editor.Function = unit.create_function("vm_set_register") func.set_address(0x1400023) editor.write("libd5A7BCF0524B8.dwarf") ``` Under the hood, LIEF uses the LLVM's DWARF backend to create and generate the final DWARF. In contrast to LLVM's low-level API, LIEF provides an abstraction that simplifies the implementation details of the DWARF format. For instance, if we want to create a DWARF for a function that contains a **stack variable** at the offset (on the stack) `8`, we can use the following API: ```python func: lief.dwarf.editor.Function = unit.create_function("vm_set_register") var: lief.dwarf.editor.Variable = func.create_stack_variable("my_stack_variable") var.set_stack_offset(8) ``` This code generates the following DWARF: ```text 0x0000000c: DW_TAG_compile_unit DW_AT_producer ("Generated by LIEF (LLVM backend)") 0x00000011: DW_TAG_subprogram DW_AT_name ("vm_set_register") DW_AT_entry_pc (0x0000000001400023) 0x0000001e: DW_TAG_variable DW_AT_name ("my_stack_variable") DW_AT_location (DW_OP_fbreg -8) ``` Defining the `DW_AT_location` for a stack variable is not as simple as it sounds. It requires defining some kind of DWARF expression and if you are curious about the actual implementation, you can check this [GitHub Gist](https://gist.github.com/romainthomas/1f7ba555c4439d9b4b457e293834ec5a). **Summary** LIEF exposes a high-level API to create DWARF based on the LLVM's low-level API ## DWARF and Reverse Engineering Reverse engineering tools typically use their own format to store information about analyzed binaries such as `*.idb` and `*.bndb`. Most of these tools are not compatible with each other, except Binary Ninja which has a support for loading IDB ([*Migrating from IDA*](https://docs.binary.ninja/guide/migration/migrationguideida.html)). For Ghidra, importing IDA database is a non-goal (c.f. [issue #2921](https://github.com/NationalSecurityAgency/ghidra/issues/2921)) One alternative is to export binary information using [BinExport](https://github.com/google/binexport) or [quokka](https://github.com/quarkslab/quokka), but many tools lack support for **importing** the exported data. In contrast, Binary Ninja, Ghidra, and IDA all have built-in support for loading DWARF files and external DWARF files. The DWARF format is primarily designed to hold information about the original source code, and the purpose of reverse engineering is to recover the semantic of the source code information from the binary. Therefore, we could use the DWARF as a reverse-engineering shared format to export types, functions, and variables from reverse-engineered binaries. ![LIEF plugins export Binary Ninja and Ghidra analysis to DWARF for use across reverse-engineering tools](https://lief.re/blog/2025-05-27-dwarf-editor/tool2dw.webp) DWARF is compatible with PE binaries, even though it is not the default format for storing debug information on Windows. **PE / DWARF** If you compile a Windows executable with `clang[-cl]` and with the flags `-g -gdwarf-5`, the final PE will contains DWARF information along with an external `.pdb`. Currently, Binary Ninja is the only tool with a built-in plugin that can generate a DWARF file from a `BinaryView` representation. However, it lacks the ability to export stack-based variables, which can be crucial information. The next section introduces two plugins for Ghidra and BinaryNinja to generate DWARF from these tools. ## BinaryNinja & Ghidra Plugins ![Binary Ninja and Ghidra plugin logos](https://lief.re/blog/2025-05-27-dwarf-editor/plugins.webp) To provide some background on this feature, I initially developed the BinaryNinja's DWARF exporter plugin for my own needs before Vector35 team released an official plugin in BinaryNinja 3.5. I use this plugin in my reverse engineering workflow to symbolize [QBDI](https://github.com/QBDI/QBDI) traces from DWARF information: 1. I statically reverse-engineer the binary 2. I generate a DWARF file 3. I trace the binary with QBDI that uses the DWARF file to symbolize: - Stack accesses (hence the need to stack variables in the DWARF) - Function calls and their parameters - Static variables accesses 4. goto 1. For instance, I used this process to reverse engineer the [DroidGuard VM](https://www.romainthomas.fr/publication/22-sstic-blackhat-droidguard-safetynet/) a few years ago. I'll take this blog post as an opportunity to share the DWARF associated with my reverse engineering of the VM `libd5A7BCF0524B8.so`: - **Original Binary:** [libd5A7BCF0524B8.so](https://lief.re/blog/2025-05-27-dwarf-editor/libd5A7BCF0524B8.so) - **Binary with generated DWARF:** [libd5A7BCF0524B8.so.debug](https://lief.re/blog/2025-05-27-dwarf-editor/libd5A7BCF0524B8.so.debug) As mentioned earlier, this functionality is integrated into BinaryNinja since version 3.5, so I'll focus more on the Ghidra plugin. For more information about the Binary Ninja integration, visit [LIEF - Binary Ninja](https://lief.re/doc/latest/plugins/binaryninja/index.html). The Ghidra plugin allows us to export Ghidra's Program information into a DWARF file. This can be done from the Project Manager interface by selecting the `DWARF` format in the export section: ![Ghidra Project Manager export dialog with DWARF selected as the output format](https://lief.re/blog/2025-05-27-dwarf-editor/project-dwarf-export.webp) You can also use this plugin from the `CodeBrowser` tool, by left-clicking on the LIEF menu and selecting `Export as DWARF`: ![Ghidra CodeBrowser LIEF menu with the Export as DWARF command](https://lief.re/blog/2025-05-27-dwarf-editor/codebrowser-export-dwarf.webp) The plugin is primarily written in Java (using the JNI) and you can also generate a DWARF file from a headless Java script: ```java import lief.ghidra.core.dwarf.export.Manager; import lief.ghidra.core.NativeBridge; public class LiefDwarfExportScript extends GhidraScript { @Override protected void run() throws Exception { NativeBridge.init(); Manager manager = new Manager(currentProgram); File output = new File("/home/romain/output.dwarf"); manager.export(output); } } ``` **IDA Support** I do not plan to support IDA for this functionality. However, if there is strong demand for it, feel free to reach out to me. You can also create your own IDA script using the Python or C++ API. ## Last Word This DWARF export functionality is still in early development, so I cannot guarantee it is free of bugs. The current version also does not export comments, but I plan to support this feature in the future. The source code for the plugins is here: - [plugins/binaryninja](https://github.com/lief-project/LIEF/tree/main/plugins/binaryninja) - [plugins/ghidra](https://github.com/lief-project/LIEF/tree/main/plugins/ghidra) Thank you for using LIEF. Romain. [^bn-dwarf-plugin]: https://github.com/Vector35/binaryninja-api/tree/f3c7443839a69bdb49b8313e6eef7a3834feb85f/plugins/dwarf/dwarf_export --- # PE Support Enhancements > PE improvements in LIEF: ARM64EC and WoW64 CHPE metadata parsing, a rebuilt PE builder for imports, resources, and TLS, and LIEF-backed ARM64EC support in Binary Ninja. - Canonical URL: https://lief.re/blog/2025-02-16-arm64ec-pe-support/ - Markdown: https://lief.re/blog/2025-02-16-arm64ec-pe-support/index.md - Authors: Romain Thomas - Published: 2025-02-16T00:00:00Z - Modified: 2025-02-16T00:00:00Z - Tags: PE, ARM64EC, Windows, Binary Ninja **I'm pleased to share that LIEF's PE support has been significantly improved on both aspects: parsing and writing.** ![Featuring Image](https://lief.re/blog/2025-02-16-arm64ec-pe-support/featured.webp) ## Parsing Improvements One of the significant updates in LIEF is the enhanced processing of the `IMAGE_LOAD_CONFIG_DIRECTORY` which now includes a detailed parsing of the underlying structures it references. For instance, the *Dynamic Value Relocation Table*[^dnv-dvrt] is now fully supported and can be accessed through the Python, C++, and Rust APIs: ```python import lief pe = lief.PE.parse("win11_arm64x_Windows.Media.Protection.PlayReady.dll") dyn_reloc_table = pe.load_configuration.dynamic_relocations for r in dyn_reloc_table: print(r) ``` ```text Dynamic Value Relocation Table (version: 1) Symbol VA: 0x0000000000000006 (RELOCATION_ARM64X) Fixup RVAs (ARM64X) [0000] RVA 0x0000010c, 2 bytes, target value 64:86 [0001] RVA 0x00000130, 4 bytes, target value d0:bc:57:00 [0002] RVA 0x00000190, 4 bytes, target value 30:22:d8:00 [0003] RVA 0x00000194, 4 bytes, target value 30:01:00:00 ... ``` All the structures involved in these dynamic relocations can be accessed through an API, including support for special relocations: - `IMAGE_DYNAMIC_RELOCATION_ARM64X` - `IMAGE_DYNAMIC_RELOCATION_FUNCTION_OVERRIDE` - `IMAGE_DYNAMIC_RELOCATION_ARM64_KERNEL_IMPORT_CALL_TRANSFER` - `IMAGE_DYNAMIC_RELOCATION_GUARD_IMPORT_CONTROL_TRANSFER` With the introduction of this new dynamic relocation support, LIEF can now offer a **relocated view** of the ARM64EC binary that is embedded within an ARM64X binary. To achieve this, simply parse the ARM64X PE binary with all options enabled. ```rust let pe = lief::pe::Binary::parse_with_config( "win11_arm64x_Windows.Media.Protection.PlayReady.dll", lief::pe::ParserConfig::with_all_options(), ).unwrap(); if let Some(nested_arm64ec) = pe.nested_pe_binary() { println!("{:?}", nested_arm64ec.header()); } ``` or in Python: ```rust pe = lief.PE.parse("win11_arm64x_Windows.Media.Protection.PlayReady.dll", lief.PE.ParserConfig.all) if nested_arm64ec := pe.nested_pe_binary: print(nested_arm64ec.header) ``` ```text Signature: 50 45 00 00 Machine: AMD64 Number of sections: 19 Pointer to symbol table: 0x0 Number of symbols: 0 Size of optional header: 0xf0 Characteristics: EXECUTABLE_IMAGE, LARGE_ADDRESS_AWARE, DLL Timtestamp: 1966221527 ``` **Parser Config** To maintain the performance of LIEF v0.17.0 in line with previous versions, the parsing of nested ARM64EC binaries must be **explicitly** enabled during the `lief.PE.parse()` operation (the default is false). LIEF provides helper functions such as: `lief.PE.Binary.is_arm64x()` and `lief.PE.Binary.is_arm64ec()` to determine whether a given PE is ARM64 emulation compatible (EC) or if it's a *fat* ARM64X. These helper functions rely on the CHPE metadata (Compiled Hybrid Portable Executable) which is accessible through LIEF for both: ARM64/ARM64EC binaries and legacy x86 binaries executed via WoW64: ```python import lief # ARM64 CHPE pe = lief.PE.parse("arm64x_ImagingEngine.dll") chpe: lief.PE.CHPEMetadataARM64 = pe.load_configuration.chpe_metadata print(chpe_metadata.auxiliary_delay_import) print(chpe_metadata.code_ranges) # x86/WoW64 CHPE pe = lief.PE.parse("Windows.Media.dll") chpe: lief.PE.CHPEMetadataX86 = pe.load_configuration.chpe_metadata print(chpe_metadata.compiler_iat_pointer) print(chpe_metadata.wowa64_dispatch_ret_function_pointer) ``` LIEF can now also process in-depth exception information for x86-64 and ARM64 binaries as well as COFF strings and COFF symbols that are still used by some toolchains. ## Writing Improvements As mentioned at the beginning of this blog post, LIEF's PE modification engine, also named Builder, has been completely refactored to enhance its reliability and consistency when modifying PE binaries. To support this update, the documentation includes a dedicated section detailing the types of modifications supported by LIEF and their limitations: - [Imports Modification](https://lief.re/doc/latest/formats/pe/modifications/imports.html) - [Resources Modification](https://lief.re/doc/latest/formats/pe/modifications/resources.html) - [TLS Modification](https://lief.re/doc/latest/formats/pe/modifications/tls.html) - [Debug Modification](https://lief.re/doc/latest/formats/pe/modifications/debug.html) - [Exports Modification](https://lief.re/doc/latest/formats/pe/modifications/exports.html) For instance, [PE import modification](https://lief.re/doc/latest/formats/pe/modifications/imports.html) has been completely redesigned to avoid using *trampolines*. ![PE Import Table for MSVC Binaries](https://lief.re/blog/2025-02-16-arm64ec-pe-support/msvc_layout.webp) Likewise, LIEF can now commit changes made to the export table. This enables the creation of new exports that can be used to expose reverse-engineered functions or to deceive disassemblers, as mentioned in [The Poor Man's Obfuscator](https://www.romainthomas.fr/publication/22-pst-the-poor-mans-obfuscator/) ![BinaryNinja Result](https://lief.re/blog/2025-02-16-arm64ec-pe-support/lief-bn-square.webp) ## BinaryNinja ARM64EC Support The ARM64EC ABI[^arm64ec], created by Microsoft, was developed to ease interoperability between ARM64 and x86-64 architectures, particularly for parts of the code that require emulation (EC = Emulation Compatible). The specifications of this ABI have a somewhat counterintuitive aspect: the targeted architecture indicated in the PE header is `AMD64`, while the majority of the functions are compiled for ARM64. Darek Mihocka, one of the authors of the ARM64EC ABI, details some of these design choices in a great blog post[^abc_arm64ec_explained] that also provides a historical background about the support of new architectures in Windows. Because the PE header specifies the x86-64 architecture, which does not accurately reflect most of the functions compiled in the binary, many disassemblers become confused when trying to analyze such a binary: ![IDA analysis of an ARM64EC binary showing many functions missing from the recognized function list](https://lief.re/blog/2025-02-16-arm64ec-pe-support/ida.webp) ![Binary Ninja ARM64EC analysis before the LIEF plugin, with only a few functions identified](https://lief.re/blog/2025-02-16-arm64ec-pe-support/bn_a.webp) To address this incomplete support, I developed a BinaryNinja plugin designed to identify additional functions and accurately determine the architecture for which each function has been compiled. For example, after using this plugin, we obtain a feature map in BinaryNinja showing that a majority of the functions have been correctly recognized: ![Binary Ninja ARM64EC analysis after the LIEF plugin, with many more functions identified](https://lief.re/blog/2025-02-16-arm64ec-pe-support/bn_b.webp) The plugin is written in Rust using both LIEF & BinaryNinja Rust bindings. This also serves as an opportunity to evaluate the LIEF Rust API and its integration into other projects. First, we can get a LIEF PE instance from a `BinaryView` with: ```rust fn enhance_with_lief(bv: &BinaryView) { let filename = bv.file().filename(); let pe = lief::pe::Binary::parse_with_config( filename.as_str(), // This is required to get access to the exceptions info that are used later lief::pe::ParserConfig::with_all_options(), ).unwrap(); } ``` **Memory Footprint** Memory-wise, parsing the PE loaded by BinaryNinja with LIEF is not ideal since the same binary exists in memory twice: once for the BinaryNinja instance and once for the LIEF instance. I think that the BinaryNinja team will enhance the support for ARM64EC binaries at some point, so in the meantime this memory overhead is acceptable. Once we have this instance, we can use the exceptions table as a source of function addresses allowing BinaryNinja to begin disassembling the binary from these addresses: ```rust let pe: &lief::pe::Binary; let bv: &BinaryView; let windows_arm64 = Platform::by_name("windows-aarch64").unwrap().to_owned(); let windows_x64 = Platform::by_name("windows-x86_64").unwrap().to_owned(); let imagebase = pe.optional_header().imagebase(); for exception in pe.exceptions() { match exception { lief::pe::RuntimeExceptionFunction::X86_64(x64) => { let addr: u64 = imagebase + (x64.rva_start() as u64); bv.add_auto_function(&windows_x64, addr); } lief::pe::RuntimeExceptionFunction::AArch64(arm64) => { let addr: u64 = imagebase + (arm64.rva_start() as u64); bv.add_auto_function(&windows_arm64, addr); } } } bv.update_analysis(); ``` This code leads to the feature map shown in the screenshot above. Although the algorithm is straightforward, it hides an important aspect of the ARM64EC ABI: The exceptions table referenced by the data directory at the index: `IMAGE_DIRECTORY_ENTRY_EXCEPTION` points to the **X86_64 table**. To access the ARM64 exceptions table we need to refer to the CHPE metadata: ```rust let loadconfig = pe.load_configuration().unwrap(); if let Some(lief::pe::CHPEMetadata::ARM64(arm64)) = loadconfig.chpe_metadata() { let arm64_exception_table_start = arm64.extra_rfe_table(); let arm64_exception_table_size = arm64.extra_rfe_table_size(); println!("{:#x} ({} bytes)", arm64_exception_table_start, arm64_exception_table_size); } ``` LIEF abstracts the processing of CHPEMetadata and provides an iterator that outputs one of the following exceptions object: - `lief::pe::exception_x64::RuntimeFunction` - `lief::pe::exception_aarch64::RuntimeFunction` You can find more details about the exception support by LIEF in the dedicated PE changelog: https://lief.re/doc/latest/changelog/pe-0-17-0.html#pe-0170-changelog The source code for the BinaryNinja plugin is available on GitHub: [lief-project/lief-binaryninja-arm64ec](https://github.com/lief-project/lief-binaryninja-arm64ec). It contains additional features, but the code should be read through the lens of a proof of concept. [^dnv-dvrt]: https://denuvosoftwaresolutions.github.io/DVRT/dvrt.html [^arm64ec]: https://learn.microsoft.com/en-us/windows/arm/arm64ec [^abc_arm64ec_explained]: http://www.emulators.com/docs/abc_arm64ec_explained.htm --- # LIEF v0.16.0 > LIEF 0.16.0: rebuilt documentation, an assembler/disassembler in LIEF Extended, dyld shared cache extraction, first mutable Rust APIs, and nanobind 2.4 bindings. - Canonical URL: https://lief.re/blog/2024-12-10-lief-0-16-0/ - Markdown: https://lief.re/blog/2024-12-10-lief-0-16-0/index.md - Authors: Romain Thomas - Published: 2024-12-10T00:00:00Z - Modified: 2024-12-10T00:00:00Z - Tags: release, LIEF Extended, assembler, disassembler, dyld-shared-cache, Rust, Python ## Documentation Documentation is an important aspect of LIEF and since the beginning of the project, I have spent a decent amount of time keeping comprehensive and intuitive documentation. Usually, I don't include documentation updates in the changelog but in this case, I thought it could be worth sharing this experience. LIEF is written in C++ with bindings for Python and Rust. Originally, the documentation was driven by languages API (isolated from each other) and generated by Sphinx with the [Breathe](https://breathe.readthedocs.io/en/latest/) extension to reference the C++ Doxygen domain. Recently, Rust landed in the arena. Compared to Python and C++, the Rust language embeds a built-in documentation engine to process and generate in-code documentation into HTML pages. Given the new Rust bindings and the Rust built-in documentation engine, two questions emerged: 1. Do we want to add (yet) another API page for Rust? 2. How do we reference Rust API in Sphinx? For the first point, I moved from a language-driven documentation structure to a functionality-driven structure. This changes how the documentation is consumed. Instead of looking for a language's format-specific API, you first choose the task, such as ELF processing or Dyld shared cache parsing. You can **then** access the language API you need. So instead of adding another Rust API reference page, the Rust API has been transparently integrated with the new layout. ![Documentation layout changes](https://lief.re/blog/2024-12-10-lief-0-16-0/api-diff.webp) The second point has been a bit more tricky to approach. With a reverse engineering background, I really value the **cross-reference feature** provided by Sphinx: ```rst blah blah blah :py:class:`lief.ELF.Binary` another blah: :cpp:class:`LIEF::ELF::Binary` ``` Python is a built-in domain supported by Sphinx and [Breathe](https://breathe.readthedocs.io/en/latest/) extension is doing the bridge between C++ Doxygen XML files and Sphinx. For Rust, there are some attempts to create a bridge but I decided to take another path. I created a Rust-sphinx domain that cross-references to the official or nightly Rust documentation. Basically with this custom domain, the following cross-references redirect to the official or nightly documentation: ```rst :rust:module:`lief::assembly` :rust:enum:`lief::assembly::Instructions` ``` Is translated into: ```text https://lief-rs.s3.fr-par.scw.cloud/doc/latest/lief/assembly/index.html https://lief-rs.s3.fr-par.scw.cloud/doc/latest/lief/assembly/enum.Instructions.html ``` By doing so, we can leverage Sphinx's cross-reference functionalities while still keeping the built-in Rust documentation. In addition to this Rust-specific domain, I created a `.. lief-api::` directive that can pack similar cross-language API into a single block. For instance, this directive: ```rst .. lief-api:: lief.Binary.disassemble() :rust:method:`lief::generic::Binary::disassemble [trait]` :rust:method:`lief::generic::Binary::disassemble_symbol [trait]` :rust:method:`lief::generic::Binary::disassemble_address [trait]` :rust:method:`lief::generic::Binary::disassemble_slice [trait]` :cpp:func:`LIEF::Binary::disassemble` :py:meth:`lief.Binary.disassemble` :py:meth:`lief.Binary.disassemble_from_bytes` ``` Is rendered as: ![Documentation layout changes](https://lief.re/blog/2024-12-10-lief-0-16-0/rust-domain.webp) This allows us to refer the API for different languages without being too verbose and impacting readability. Combined with Sphinx substitution, we can write: ``` This is an example that cross-reference |lief-disassemble| .. |lief-disassemble| lief-api:: lief.Binary.disassemble() :rust:method:`lief::generic::Binary::disassemble [trait]` :rust:method:`lief::generic::Binary::disassemble_symbol [trait]` :rust:method:`lief::generic::Binary::disassemble_address [trait]` :rust:method:`lief::generic::Binary::disassemble_slice [trait]` :cpp:func:`LIEF::Binary::disassemble` :py:meth:`lief.Binary.disassemble` :py:meth:`lief.Binary.disassemble_from_bytes` ``` You can go checking out this page [https://lief.re/doc/latest/formats/pe/index.html](https://lief.re/doc/latest/formats/pe/index.html) to see a concrete rendering of these changes. ## Extended Features **Public Release** The extended version is now publicly available at this address: [https://extended.lief.re](https://extended.lief.re) ## Assembler & Disassembler Adding (or not adding) a disassembler in LIEF has been a long-standing question and with the [extended](https://lief.re/doc/latest/extended/intro.html) version, I found a fair trade-off: ** LIEF core focuses on executable formats, free from any extra features that might have a significant impact on the build complexity or library size. ** On the other hand, LIEF extended provides additional functionalities that require a more complex build pipeline and increase the binary size. Among these extended functionalities, there are a [disassembler](https://lief.re/doc/latest/extended/disassembler/index.html) and an [assembler](https://lief.re/doc/latest/extended/assembler/index.html) based on the LLVM's MC layer. The disassembling API is provided at different levels: ### LIEF::Binary ```python import lief pe = lief.PE.parse("cmd.exe") for inst in pe.disassemble(0x400000): print(inst) # Instruction semantic print(inst.is_syscall) print(inst.is_memory_access) print(inst.is_call) # Instruction operands (for AArch64 and x86-64) if isinstance(inst, lief.assembly.aarch64.Instruction): for idx, operand in enumerate(inst.operands): match operand: case lief.assembly.aarch64.operands.Register(): print(f"OP[{idx}] -- REG: {operand.value}") case lief.assembly.aarch64.operands.Memory(): print(f"OP[{idx}] -- MEM: {operand.base} {operand.offset}") case lief.assembly.aarch64.operands.PCRelative(): print(f"OP[{idx}] -- PCR: {operand.value}") case lief.assembly.aarch64.operands.Immediate(): print(f"OP[{idx}] -- IMM: {operand.value}") ``` ### LIEF::dwarf::Function ```python import lief elf = lief.ELF.parse("my-dbg.elf") dwarf: lief.dwarf.DebugInfo = elf.debug_info func: lief.dwarf.Function = dwarf.find_function("my_debug_function") for inst in func.instructions: print(inst) ``` ### LIEF::dsc::DyldSharedCache ```python import lief cache = lief.dsc.load("ios-18/") for inst in cache.disassemble(0x1886f4a44): print(inst) ``` In terms of implementation, the disassembler wraps a lazy iterator that evaluates/disassembles an instruction **only** when the iterator is processed. It means that you don't pay any overhead until you access the iterator's value: ```text # O(0) inst = macho.disassemble(0x400000) inst = macho.disassemble(0x400000) # O(10) for _ in range(10): next(inst) ``` The `.end()` sentinel of the iterator is based on two properties: 1. Either a range is specified (e.g. `macho.disassemble(0x400000, /*size*/0x1000)`) and the iterator past the end of the range. 2. The instruction can't be disassembled. This kind of sentinel allows us to use this API: `macho.disassemble(0x400000)` which will disassemble (lazily) instructions at the address `0x400000` until it fails. **C++ & Rust & Python** The disassembler/assembler API is uniformly available in Rust, C++, and Python. ### Capstone? Nyxstone? [As stated in the documentation](https://lief.re/doc/latest/extended/disassembler/index.html#technical-details) the major design difference with [Capstone](https://www.capstone-engine.org/) is that LIEF uses a mainstream version of LLVM with limited patches[^llvm-patch] on the MC layer (the current version is based on LLVM `19.1.2`). The design difference with [Nyxstone](https://github.com/emproof-com/nyxstone) is that LLVM is hidden from the public API which means that it does not require to have an LLVM version pre-install on the system. Moreover, it exposes opcodes and control-flow/semantic information about the instructions. **On the other hand, LIEF does not provide a standalone API to disassemble arbitrary instructions. The disassembler engine is bound to the object from which the API is exposed.** ### Assembler In association with a disassembler, LIEF exposes a (basic) assembly API that allows generating **and patching** instructions: ```python import lief elf = lief.ELF.parse("my-android-obfuscated.so") text = elf.get_section(".text") # Disassembler syscall = [inst for inst in elf.disassemble(bytes(text)) if inst.is_syscall] # Assembler for syscall_inst in syscall: new_bytes = elf.assemble(syscall_inst.address, "nop;") # Assemble AND patch print(new_bytes.hex(", ")) ``` **Warning** In this current version, the assembler is working *pretty* well for x86/x86_64 and AArch64 but might break on other architectures. In addition, `llvm::MCFixup`** are not supported.** This can be used to patch LIEF's binary object directly at the assembly level. I plan to provide the assembly engine with LIEF Binary context. If the binary defines a function such as `call_me()` that is exported or present in the debug information, users would be able to call it at the assembly level: ```rust fn patch_with_context(macho: &mut lief::macho::Binary) { macho.assemble(0x140000090, r#" adrp x0, call_me; add x0, x0, :lo12:call_me; mov x1, 0x90; str x1, [x0]; "#r); } ``` And LIEF would handle the relocation/resolution process to instruct LLVM about the location and the definition of `call_me`. **C++ & Rust & Python** The disassembler/assembler API is seamlessly available in Rust, C++, and Python :) ## Dyld Shared Cache Initial support for processing Apple's Dyld shared cache with LIEF has been released along with an API to **deoptimize** in-cache Dylib. The API looks like this: ```python import lief cache = lief.dsc.load("ios-18.1/") for dylib in cache.libraries: print(f"0x{dylib.address:016x} {dylib.path}") # Extract the dylib as a regular lief.MachO.Binary macho: lief.MachO.Binary = dylib.get() ``` **Warning** Please note that the deoptimization feature is not working well on all the shared cache libraries. This support is going to be improved over time. One could also use this API to diff two shared caches: ```rust use lief; let ios_17 = lief::dsc::load_from_path("ios-17.7.1"); let ios_18 = lief::dsc::load_from_path("ios-18.1.1"); let libraries_17: HashSet = ios_17.libraries() .map(|lib| lib.path()) .collect(); let libraries_18: HashSet = ios_18.libraries() .map(|lib| lib.path()) .collect(); println!("{:?}", libraries_17.symmetric_difference(&libraries_18)) ``` ## Rust Rust bindings got their first **mutable** functions which are listed in the [changelog](https://lief.re/doc/latest/changelog.html). These mutable functions are limited but they allow us to make basic modifications like adding a library or patching assembly code: ```rust fn add_library(elf: &mut lief::elf::Binary) { elf.add_library("libtest.so"); elf.write("patched.elf"); } ``` ```rust fn patch_asm(elf: &mut lief::macho::Binary) { macho.assemble(0x100004090, r#" mov x0, x16; br x0; "#); macho.write("patched.macho"); } ``` In addition, the support for the `x86_64-unknown-linux-musl` target triple is now available and the minimal GLIBC version for `x86_64-unknown-linux-gnu` has been lowered to `2.28`. It means that Linux Rust bindings can now run on Debian 10, Ubuntu 19.10, ... while before it required Debian 11 or Ubuntu 20.04. The new `x86_64-unknown-linux-musl` triple can be used to generate **full static** without **any dependencies** to the `libstdc++, libc, ...`. For instance, given this code: ```rust use lief; use lief::generic::Section; fn main() { let path = std::env::args().last().unwrap(); let mut file = std::fs::File::open(path).expect("Can't open the file"); if let Some(lief::Binary::PE(pe)) = lief::Binary::from(&mut file) { for section in pe.sections() { println!( "{:20}: [0x{:016x}-0x{:016x}]", section.name(), section.virtual_address(), section.virtual_address() + section.virtual_size() as u64 ); } } } ``` We can generate a dependencies-free executable by running: ```console $ cargo build [--release] --target x86_64-unknown-linux-musl ``` ```console $ ldd target/x86_64-unknown-linux-musl/release/reader statically linked ``` ```console $ target/x86_64-unknown-linux-musl/release/reader steam.exe .text : [0x0000000000001000-0x00000000002cbe53] .rdata : [0x00000000002cc000-0x00000000003a7fa2] .data : [0x00000000003a8000-0x000000000043ada0] .rsrc : [0x000000000043b000-0x0000000000471b8c] .reloc : [0x0000000000472000-0x0000000000490c74] ``` ## Python Bindings LIEF is now using [nanobind](https://nanobind.readthedocs.io/en/latest/) v2.4.0 which improves the support for typing. Among these typing improvements, C++ enums flags are now properly inheriting from `enum.Flag` which results in a better interface with Python code. Typing stub files (`*.pyi`) are now also generated with nanobind's built-in `stubgen.py` instead for [mypy](https://github.com/python/mypy/blob/ac8957755a35a255f638c122e22c03b0e75b9a79/mypy/stubgen.py). ## Final Words Additional changes are listed in the detailed [changelog](https://lief.re/doc/latest/changelog.html). Many thanks to dornstetter and [kohnakagawa](https://github.com/kohnakagawa) for their feedback about the dyld shared cache feature. Thank you also to [Konstantin Vinogradov](https://github.com/vinogradovkonst) and [dctoralves](https://github.com/mateeuslinno) for their sponsorship. [^llvm-patch]: All the patches have been PR-submitted to the LLVM. You can check [LIEF & LLVM](https://lief.re/doc/latest/extended/intro.html#lief-extended-llvm) for the details --- # LIEF v0.15.0 > LIEF 0.15.0 introduces official Rust bindings and LIEF Extended, with Objective-C, DWARF, PDB, parser performance, and Python wheel updates. - Canonical URL: https://lief.re/blog/2024-07-21-lief-0.15-0/ - Markdown: https://lief.re/blog/2024-07-21-lief-0.15-0/index.md - Authors: Romain Thomas - Published: 2024-07-21T00:00:00Z - Modified: 2024-07-21T00:00:00Z - Tags: release, Rust, LIEF Extended, Objective-C, DWARF, PDB, Python While this new release adds new functionalities and addresses different bugs, It is worth mentioning that it is the first release to officially expose Rust binding! In addition, an *extended* version was also released to provide additional functionalities not strictly related to the executable formats. ## Rust bindings As discussed in these blog posts: 1. [LIEF Rust bindings updates](https://lief.re/blog/2024-06-16-rust-update/) 2. [Rust bindings for LIEF](https://lief.re/blog/2024-04-28-rust/) LIEF is now available in Rust for the following architectures: - `aarch64-unknown-linux-gnu` - `x86_64-apple-darwin` - `x86_64-pc-windows-msvc` (MT/MD runtimes) - `x86_64-unknown-linux-gnu` - `aarch64-apple-ios` - `aarch64-apple-darwin` I published the release on [crates.io](https://crates.io/crates/lief) so you should be able to start using LIEF in Rust with: ```rust [package] name = "lief-demo" version = "0.0.1" edition = "2021" [dependencies] lief = "0.15.0" ``` ## LIEF Extended LIEF is now providing additional features thanks to an extended version. Among those features, it provides support for DWARF and PDB debug formats as well as Objective-C metadata. ## Objective-C This support is a kind of spin-off of [iCDump](https://www.romainthomas.fr/post/23-01-icdump/) which is now completely integrated into LIEF. Compared to the original [iCDump](https://github.com/romainthomas/iCDump) project, it fixes the issue with the new chained relocations (c.f. [issue#4](https://github.com/romainthomas/iCDump/issues/4)) format and can be used on all the platforms supported by LIEF (including Windows) in C++/Rust/Python: **Rust:** ```rust let macho: lief::macho::Binary; if let Some(metadata) = macho.objc_metadata() { println!("Objective-C metadata found"); for class in metadata.classes() { println!("name={}", class.name()); for method in class.methods() { println!(" method.name={}", method.name()); } } } ``` **Python:** ```python import lief macho: lief.MachO.Binary = ... metadata: lief.objc.Metadata = macho.objc_metadata if metadata is not None: print("Objective-C metadata found") for clazz in metadata.classes: print(f"name={clazz.name}") for meth in clazz.methods: print(f" method.name={meth.name}") # Generate a header like "class-dump" print(metadata.to_decl()) ``` ## DWARF & PDB ![DWARF & PDB Hierarchy](https://lief.re/blog/2024-07-21-lief-0.15-0/dwarf-pdb-hierarchy.webp) Supporting debug formats like DWARF or PDB has been a long-standing discussion (c.f. [issue #17](https://github.com/lief-project/LIEF/issues/17)). The main reasons to avoid supporting these formats from scratch were: 1. The maintenance effort 2. There already exists libraries to process these debug formats: - [pyelftools](https://github.com/eliben/pyelftools) for DWARF - [LLVM](https://llvm.org/) (DWARF & PDB) - [gimli](https://docs.rs/gimli/latest/gimli/) (DWARF) On the other hand, I understand the need to process debug information from a LIEF binary object when it is present. The existing projects expose powerful low-level APIs that match the debug format specifications. However, they do not provide[^sense] an abstraction over the complexity of these specifications. Developers and reverse engineers work with concepts such as compilation units, functions, global variables, and stack variables. Before accessing this information from a DWARF or PDB file, however, you need to understand a PDB DBI stream or know that a function's address in DWARF can be determined by either `DW_AT_entry_pc` or `DW_AT_low_pc`. The idea behind the support of the DWARF and PDB formats in LIEF is to: 1. **bridge concepts that make sense to the developers/reverse engineers with their concrete specifications in DWARF/PDB** 2. Have a (documented) C++ API and bindings for Python/Rust. This LIEF bridge is **based on LLVM** which did the heavy job of supporting DWARF & PDB within a single framework. The DWARF & PDB support in LIEF leverages the LLVM API to abstract concepts as listed above. For instance, you can iterate through every public symbol in the `ntoskrnl.pdb` PDB through: ```python import lief ntoskrnl: lief.pdb.DebugInfo = lief.pdb.load("./ntoskrnl.pdb") for sym in ntoskrnl.public_symbols: print(f"{sym.demangled_name}: 0x{sym.RVA:06x}") ``` If the PDB embeds extended information about the compilation units we can do (in Rust): ```rust let pdb = lief::pdb::load("peacecannary.pdb"); for cu in pdb.compilation_units() { for func in cu.functions() { if func.name().starts_with("peacecannary::CObfuscator") { println!("{}: {} (0x{:04x})", cu.module_name(), func.name(), func.rva()); } } } ``` The API for the DWARF format is pretty similar: ```python import lief elf: lief.ELF.Binary = ... # If the binary embeds DWARF debug info in the ELF: dwarf: lief.dwarf.DebugInfo = elf.debug_info # Otherwise: dwarf: lief.dwarf.DebugInfo = lief.dwarf.load("my_dwarf.dwarf") for cu in dwarf.compilation_units: print(f"Produced by: {cu.producer} in {cu.compilation_dir}") for func in cu.functions: print(f"0x{func.address:04x}: {func.name} ({func.size} bytes)") for var in cu.variables: if var.is_constexpr: continue # Look for global variables only if var.address is not None and var.address > 0: print(f"0x{var.address:04x}: {var.linkage_name} ({var.size} bytes)") ``` For more details about the API, you can take a look at these dedicated sections: - [DWARF](https://lief.re/doc/latest/extended/dwarf/) - [PDB](https://lief.re/doc/latest/extended/pdb/) [^sense]: Which makes sense since this is not the purpose of these projects ## Other Updates ## Mach-O AI LIEF is now ~powered by AI~ supporting Apple `*.hwx` files which are some kind of Mach-O file for the Apple Neural Engine (ANE). These `*.hwx` start with a new magic identifier: `0xbeefface` and embed custom `LC_` command like the command `0x40` **LC Command 0x40** I could be interested in adding the support of this *private* command in LIEF so if anyone already reversed or has some info about the layout of this command, feel free to reach out. To support unknown or non-public LC commands in LIEF, I created an artificial `LIEF::MachO::UnknownCommand` which is a placeholder for any Mach-O commands that are not recognized by LIEF. For instance, we can inspect the private `0x40` command as follows: ```python import lief target = lief.MachO.parse("personsemantics-u8-v4.H16.espresso.hwx").at(0) lc_0x40: lief.MachO.UnknownCommand = macho.commands[18].command print(lc_0x40.original_command) # Outputs 0x40/61 print(bytes(lc_0x40.data)) # Print the raw content of the command ``` These `.hwx` files have been involved in the [Dopamine jailbreak](https://github.com/opa334/Dopamine/blob/43c03c167ccaa23ca51f268213e5abc85a9aee55/Application/Dopamine/Exploits/weightBufs/exploit/exploit.m#L514) and you can also find a BlackHat presentation about the Apple Neural Engine: [Apple Neural Engine Internal](https://i.blackhat.com/asia-21/Friday-Handouts/as21-Wu-Apple-Neural_Engine.pdf). ## PE Authenticode LIEF can inspect and verify the PE Authenticode and with this release, we can even do that in Rust! ```rust use lief::pe; let mut file = std::fs::File::open(path).expect("Can't open the file"); if let Some(lief::Binary::PE(pe)) = lief::Binary::from(&mut file) { let result = pe.verify_signature(pe::signature::VerificationChecks::DEFAULT); if result.is_ok() { println!("Valid signature!"); } else { println!("Signature not valid: {}", result); } return ExitCode::SUCCESS; } ``` This new release also adds support for the MS-CounterSignature attribute (OID: `1.3.6.1.4.1.311.3.3.1`) and some other attributes like `Ms-ManifestBinaryID` (OID: `1.3.6.1.4.1.311.10.3.28`) ## ELF No breaking updates for the ELF format. LIEF is now able to parse and modify binaries compiled with the new `DT_RELR` and `DT_ANDROID_REL_` relocations. ![ELF Dynamic Array Relocated](https://lief.re/blog/2024-07-21-lief-0.15-0/elf-relocations.webp) I also added the helper: `LIEF::ELF::Binary::get_relocated_dynamic_array` which allows us to get a *relocated* view of the `DT_INIT_ARRAY/DT_FINI_ARRAY`. This can be useful when -- for instance -- the init array values are null because of relocations: ```python import lief elf: lief.ELF.Binary = ... # Return: [0, 0, 0, 0, ...] elf.get(lief.ELF.DynamicEntry.TAG.INIT_ARRAY).array # Return relocated values: [0x96db10, 0x9b9c14, 0xe7f660, 0xe7f70c, ...] elf.get_relocated_dynamic_array(lief.ELF.DynamicEntry.TAG.INIT_ARRAY) ``` ## Enums Since the beginning of LIEF, all the enums used by the different formats were located in a **single** header file (e.g. `LIEF/PE/enums.hpp` or `lief.PE.{enums, ...}` in Python). Some of them were clashing with system headers that were also `#define` some of these enums. To work around this issue, we had a dirty hack based on `LIEF/{ELF.PE,MachO}/undef.h` that undefines these values before being included. In LIEF 0.15.0 the scope of the enums has been redefined so that we should no longer need the `undef.h`. For instance the standalone enum `LIEF::ELF::ELF_SECTION_TYPES` (or `lief.ELF.SECTION_TYPES`) has been re-scoped in the `LIEF::ELF::Section` class: ```cpp // class LIEF_API Section : public LIEF::Section { enum class TYPE : uint64_t { SHT_NULL = 0, /**< No associated section (inactive entry). */ PROGBITS = 1, /**< Program-defined contents. */ ... }; }; ``` This means that instead of using `LIEF::ELF::ELF_SECTION_TYPES::SHT_PROGBITS` or `lief.ELF.SECTION_TYPES.SHT_PROGBITS` you should now use: ```diff {style=pastie} - LIEF::ELF::ELF_SECTION_TYPES::SHT_PROGBITS + LIEF::ELF::Section::TYPE::PROGBITS - lief.ELF.SECTION_TYPES.SHT_PROGBITS + lief.ELF.Section.TYPE.PROGBITS ``` The list of the enums affected by this change is listed in the [changelog](https://lief.re/doc/latest/changelog.html). ## Performances ### PE Parser I received some feedback about performance issues in the latest release (`0.14.x`) compared to former releases. This regression affects Mach-O and PE binaries and I'm happy to say that this `v0.15.0` release should be faster on ELF, PE, and Mach-O compared to previous releases. The PE regression comes from the `LIEF::PE::OptionalHeader::computed_checksum` introduced in LIEF 0.12.0 and discussed in this issue: [#660](https://github.com/lief-project/LIEF/issues/660). As of LIEF 0.12.0, this `computed_checksum` was computed during the parsing phase, and on large binaries, this computation might have a significant impact on the performances. In LIEF 0.15.0, the OptionalHeader's checksum can be re-computed over the `LIEF::PE::Binary` object: ```python import lief pe: lief.PE.Binary = ... computed_checksum = pe.compute_checksum() ``` Thus, avoiding the computation during the parsing phase and moving to an "*on-demand*" API. ### Mach-O Parser On the other hand, the Mach-O regression was pretty tricky to identify (c.f. [issue #1069](https://github.com/lief-project/LIEF/issues/1069)). The root cause of the regression was these lines: ```cpp // https://github.com/lief-project/LIEF/blob/0.14.1/src/MachO/BinaryParser.cpp#L285-L290 for (LARGE_LOOP) { if (!is_printable(name)) { ... } } ``` with `is_printable` implemented as follows: ```cpp bool is_printable(const std::string& str) { return std::all_of(std::begin(str), std::end(str), [] (char c) { return std::isprint(c, std::locale("C")); }); } ``` Then, while processing large Mach-O binaries with LIEF we can observe: - On Linux: No regression - On macOS: **REGRESSION** - On Windows: **REGRESSION** It turned out that `std::locale("C")` is *cached* by the STL on Linux but not on macOS & Windows. This means that we were invoking `std::locale("C")` **for each character of each string** (which has a cost). One solution is to store `std::locale("C")` in a static variable as it is done -- under the hood -- in the Linux STL. ```diff {style=pastie} bool is_printable(const std::string& str) { return std::all_of(std::begin(str), std::end(str), - [] (char c) { return std::isprint(c, std::locale("C")); }); + [] (char c) { + static std::locale LC("C"); + return std::isprint(c, LC); + }); } ``` This actual fix is slightly different though: [7c3f63194](https://github.com/lief-project/LIEF/commit/7c3f631948051c8afe5c00e49635385c7de857ba). ## Python Wheels LIEF Python wheels are now available for musl-based systems. This support is motivated by the fact that Python Docker images tagged with the suffix the `-alpine` are using Alpine system which is based on musl libc. Thus, we can now use Docker's python-alpine as image base to install LIEF: ```docker FROM python:3.13.0b3-alpine RUN pip install --no-cache-dir lief==0.15.0 ``` Note that the LIEF Python wheel for Alpine weighs about 2.5 megabytes compressed and 7 megabytes decompressed. ## Final Words This new Rust-oriented release is a major milestone for LIEF. While the library is widely used among Python community with ~16,000 daily downloads on PyPI, I'm eager to see new use cases or issues brought by the Rust community. As a reminder, there is a [Discord](https://discord.gg/jGQtyAYChJ) channel where you can drop your questions, and remarks (that are not [issues](https://github.com/lief-project/LIEF/issues/new/choose) :wink:). Thank you also to [arttson](https://github.com/arttson) and [lexika979](https://github.com/lexika979), for their sponsorship. --- # LIEF Rust bindings updates > LIEF Rust bindings ahead of 0.15.0: new documentation, more supported architectures, Rust tier 1 and tier 2 coverage, and use cases such as Authenticode checks. - Canonical URL: https://lief.re/blog/2024-06-16-rust-update/ - Markdown: https://lief.re/blog/2024-06-16-rust-update/index.md - Authors: Romain Thomas - Published: 2024-06-16T00:00:00Z - Modified: 2024-06-16T00:00:00Z - Tags: Rust, bindings The rust bindings for LIEF are getting more and more production-ready for the next official release of LIEF (`v0.15.0`). This blog post exposes the recent updates on these bindings and some use cases. ## Documentation The Rust documentation for the current bindings is now almost complete such as most of the functions and structures are documented: [![LIEF Rust Documentation](https://lief.re/blog/2024-06-16-rust-update/doc.png)](https://lief-rs.s3.fr-par.scw.cloud/doc/latest/lief/pe/struct.Signature.html#method.check) One can access the nightly doc at this address: https://lief-rs.s3.fr-par.scw.cloud/doc/latest/lief/index.html ## New Architectures Supported As mentioned in the [previous blog post](https://lief.re/blog/2024-04-28-rust/), the Rust bindings work with a "pre-compilation" step. Since the previous blog post, I added the support of iOS (i.e. `aarch64-apple-ios`) and for Linux ARM64 (i.e. `aarch64-unknown-linux-gnu`) which gives us this support in LIEF compared to the [Rust Platform Support](https://doc.rust-lang.org/nightly/rustc/platform-support.html#platform-support) ### Rust Tier 1 Support | Triplet | Support | Comment | |-----------------------------|--------------------|------------------------------| | `aarch64-unknown-linux-gnu` | :white_check_mark: | | | `i686-pc-windows-gnu` | :x: | | | `i686-pc-windows-msvc` | :shrug: | Could be supported if needed | | `i686-unknown-linux-gnu` | :shrug: | Could be supported if needed | | `x86_64-apple-darwin` | :white_check_mark: | | | `x86_64-pc-windows-gnu` | :x: | | | `x86_64-pc-windows-msvc` | :white_check_mark: | | | `x86_64-unknown-linux-gnu` | :white_check_mark: | | ### Rust Tier 2 Support | Triplet | Support | Comment | |-----------------------------|--------------------|------------------------------| | `aarch64-apple-ios` | :white_check_mark: | | | `aarch64-apple-ios-sim` | :shrug: | Could be supported if needed | | `aarch64-linux-android` | :stopwatch: | Planned | | `aarch64-apple-darwin` | :white_check_mark: | | | `x86_64-unknown-linux-musl` | :stopwatch: | Planned | The support for some triplets like `i686-pc-windows-msvc` will be done on an as-needed basis so feel free to reach out or to open an issue/discussion on GitHub if you need this support. ## Uses Cases ```rust // This code checks the PE Authenticode let path = std::env::args().last().unwrap(); let mut file = std::fs::File::open(path).expect("Can't open the file"); if let Some(lief::Binary::PE(pe)) = lief::Binary::from(&mut file) { let result = pe.verify_signature(pe::signature::VerificationChecks::DEFAULT); if result.is_ok() { println!("Valid signature!"); } else { println!("Signature not valid: {}", result); } return ExitCode::SUCCESS; } ExitCode::FAILURE ``` --- ```rust // This code list all the libraries needed by an ELF binary as well as // the versioning of the symbols. // Example of output: // Dependencies: // - libclang-cpp.so.17 // - libLLVM-17.so // - libstdc++.so.6 // - libc.so.6 // Versions: // From libc.so.6 // - GLIBC_ABI_DT_RELR // - GLIBC_2.14 // - GLIBC_2.34 // - GLIBC_2.32 // From libstdc++.so.6 // - GLIBCXX_3.4.29 // - GLIBCXX_3.4.30 // From libLLVM-17.so // - LLVM_17 let mut args = std::env::args(); if args.len() != 2 { println!("Usage: {} ", args.next().unwrap()); return ExitCode::FAILURE; } let path = std::env::args().last().unwrap(); let mut file = std::fs::File::open(&path).expect("Can't open the file"); if let Some(lief::Binary::ELF(elf)) = lief::Binary::from(&mut file) { println!("Dependencies:"); for entry in elf.dynamic_entries() { if let dynamic::Entries::Library(lib) = entry { println!(" - {}", lib.name()); } } println!("Versions:"); for version in elf.symbols_version_requirement() { println!(" From {}", version.name()); for aux in version.auxiliary_symbols() { println!(" - {}", aux.name()); } } return ExitCode::SUCCESS; } println!("Can't process {}", path); ExitCode::FAILURE ``` --- ```rust // Inspecting the PE rich header let path = std::env::args().last().unwrap(); let mut file = std::fs::File::open(&path).expect("Can't open the file"); if let Some(lief::Binary::PE(pe)) = lief::Binary::from(&mut file) { let rich_header = pe.rich_header().unwrap_or_else(|| { println!("Rich header not found!"); process::exit(0); }); println!("Rich header key: 0x{:x}", rich_header.key()); for entry in rich_header.entries() { println!("id: 0x{:04x} build_id: 0x{:04x} count: #{}", entry.id(), entry.build_id(), entry.count()); } return ExitCode::SUCCESS; } println!("Can't process {}", path); ExitCode::FAILURE ``` --- ```rust // Dumping which section of an iOS app is encrypted let path = std::env::args().last().unwrap(); let mut file = std::fs::File::open(&path).expect("Can't open the file"); if let Some(lief::Binary::MachO(fat)) = lief::Binary::from(&mut file) { for macho in fat.iter() { for cmd in macho.commands() { if let lief::macho::Commands::EncryptionInfo(info) = cmd { println!("Encrypted area: 0x{:08x} - 0x{:08x} (id: {})", info.crypt_offset(), info.crypt_offset() + info.crypt_size(), info.crypt_id() ) } } } return ExitCode::SUCCESS; } ``` --- # Rust bindings for LIEF > An introduction to LIEF's Rust bindings, their idiomatic API design, memory-safety model, cross-compilation support, and engineering tradeoffs. - Canonical URL: https://lief.re/blog/2024-04-28-rust/ - Markdown: https://lief.re/blog/2024-04-28-rust/index.md - Authors: Romain Thomas - Published: 2024-04-28T00:00:00Z - Modified: 2024-04-28T00:00:00Z - Tags: Rust, bindings, memory-safety LIEF Rust bindings are now available. This blog post introduces these bindings and the technical challenges behind this journey. ## tl;dr ```toml [package] name = "lief-demo" version = "0.0.1" edition = "2021" [dependencies] lief = { git = "https://github.com/lief-project/LIEF", branch = "main"} ``` ```rust use lief::Binary; fn main() { let mut file = File::open(path).expect("Can't open the file"); match Binary::from(&mut file) { Some(Binary::ELF(elf)) => { for section in elf.sections() { println!("{}: 0x{:x}", section.name(), section.virtual_address()); } }, Some(Binary::PE(pe)) => { // ... }, Some(Binary::MachO(macho)) => { // ... }, None => { // Parsing error } } } ``` Nightly documentation is available here: https://lief-rs.s3.fr-par.scw.cloud/doc/latest/lief/index.html and the package will be published on https://crates.io/crates/lief for the `0.15.0` release. ## Introduction It has been a long journey to have Rust bindings for LIEF, and I'm happy to announce that these bindings are starting to be ready for public release. I'll take this blog post as an opportunity to share the different challenges that led me to the current design of the bindings. I'm not a Rust guru, so feel free to share your feedback or suggestions! ## Idiomatic bindings First off, I'm attached to have bindings that are idiomatic in the language they target. Reaching the current state of the Rust API took me most of the time during the development. The Rust language introduces new concepts that do not exactly match what we can find in object-oriented languages. You can get an idea of the Rust API with these examples. ### Iterate over ELF sections ```rust use lief::Binary; use lief::generic::Section; // for the "abstract" traits let path = std::env::args().last().unwrap(); if let Some(Binary::ELF(elf)) = Binary::parse(path.as_str()) { for section in elf.sections() { println!("{}", section.name()); } } ``` ### Get PE PDB path ```rust use lief::Binary; use lief::pe::debug::Entries::CodeViewPDB; if let Some(Binary::PE(pe)) = Binary::parse(path.as_str()) { for entry in pe.debug() { if let CodeViewPDB(pdb_view) = entry { println!("{}", pdb_view.filename()); } } } ``` ### Access Mach-O Dyld Info ```rust use lief::Binary; use lief::macho::commands::Commands; use lief::macho::binding_info::BindingInfo; if let Some(Binary::MachO(fat)) = Binary::parse(path.as_str()) { for macho in fat.iter() { // First version, iterate over the commands for cmd in macho.commands() { // Alternative to `if let` pattern match cmd { Commands::DyldInfo(dyld_info) => { for binding in dyld_info.bindings() { if let BindingInfo::Chained(chained) = binding { println!("Library: 0x{:x}", chained.address()); } } } _ => {} } } // Second version, using the helper if let Some(dyld_info) = macho.dyld_info() { for binding in dyld_info.bindings() { if let BindingInfo::Chained(chained) = binding { println!("Library: 0x{:x}", chained.address()); } } } } } ``` Given this idiomatic goal, there were some challenges in exposing C++ code to Rust. ### Polymorphism & Inheritance How to idiomatically bind this C++ code in Rust? ```cpp class Base { virtual std::string get_name() { return "Base"; } }; class Derived : public Base { virtual std::string get_name() { return "Derived"; } }; class OtherDerived : public Base { virtual std::string get_name() { return "OtherDerived"; } }; ``` For the inheritance relationship, the idea is to leverage Rust's `enum` structure in which, all the **leaves** of the inheritance tree are an entry of the enum: ```rust pub enum Inheritance { Derived(Derived), OtherDerived(OtherDerived), } ``` Secondly, all these objects inherit and share the `get_name()` virtual function. To provide this shared *property* in Rust, we can leverage a Rust `trait` that would make `get_name` available for the structures that implement this trait: ```rust pub trait AsBase { fn get_name(&self) -> String; } impl AsBase for Derived { fn get_name(&self) -> String { ... } } impl AsBase for OtherDerived { fn get_name(&self) -> String { ... } } ``` One can also simplify the definition of the trait such as the derived objects only have to provide the `FFI` reference to the base class: ```diff {style=pastie} pub trait AsBase { - fn get_name(&self) -> String; + fn as_base(&self) -> ffi::BaseImpl; + + fn get_name(&self) -> String { + self.as_base().get_name().to_string() + } } impl AsBase for Derived { - fn get_name(&self) -> String { + fn as_base(&self) -> ffi::BaseImpl { ... } } impl AsBase for OtherDerived { fn get_name(&self) -> String { ... } } ``` LIEF's Rust bindings highly rely on these patterns to expose classes with polymorphism and inheritance properties. ### Lifetime In C++, we don't have the concept of a lifetime for an object. For instance, it's perfectly fine to write this code: ```cpp int main() { LIEF::PE::Binary* pe = nullptr; { std::unique_ptr pe_unique = LIEF::PE::Parser::parse("..."); pe = pe_unique.get(); } printf("%s\n", pe->get_section(".text").name()); // Use-after-free return 0; } ``` Nevertheless, the `pe` pointer used in `printf` is no longer valid because of the scope of the `std::unique_ptr`. In Python, nanobind and pybind11 provide helpers to define the lifetime of an object according to its parent or its scope: ```cpp nb::class(m, "Binary") .def_prop_ro("sections", nb::overload_cast<>(&Binary::sections), nb::keep_alive<0, 1>()) ``` With `nb::keep_alive`, we indicate that the lifetime of the PE section iterator must be at least as long as the lifetime of the PE Binary instance. In Rust, we could express this lifetime with something like: ```rust pub struct Iterator<'a> { pub it: ffi::Impl, } impl<'a> Iterator<'a> { pub fn new(it: ffi::Impl) -> Self { Self { it, } } } impl Binary { pub fn get_iterator(&'a self) { Iterator::new(self.get_ffi_impl()) } } ``` But this code is not correct since the lifetime `<'a>` of the `Iterator` structure is not bound to an attribute in the structure. For technical-ffi reasons, we can't bind this lifetime to `ffi::Impl`. One solution consists of using [PhantomData](https://doc.rust-lang.org/std/marker/struct.PhantomData.html) to provide the lifetime semantic: ```rust pub struct Iterator<'a> { pub it: ffi::Impl, _owner: PhantomData<&'a ffi::PE_Binary>, } ``` ## Safety First! LIEF is developed in what we could say, an "unsafe" language (i.e. C++). On the other hand, Rust provides strong guarantees about memory, concurrency, ... Even though LIEF's core can't provide the safety guarantees that Rust is giving, I tried to provide *some* guarantees about the bindings. ### Coverage 65% of the functions exposed by the Rust binding are covered by the test suite and you can access the coverage report here: https://lief-rs.s3.fr-par.scw.cloud/coverage/index.html (nightly generated). **Coverage** By covered I mean: *"the function that bridges from C++ to Rust **is executed** in the test suite"*. ### ASAN Regarding memory safety, Rust allows packages to be compiled with ASAN through compiler options: ``` export RUSTFLAGS="-Z sanitizer=address -Clink-args=-fsanitize=address" export TARGET_CXXFLAGS="-fsanitize=address -fno-omit-frame-pointer -O1" ... ``` Thus, we also leverage this option to compile **both**: LIEF core and the Rust binding with ASAN. Given the fact that 65% of the functions and 70% of the lines are test-covered, running these tests with ASAN gives us some confidence about the fact that the bindings do not introduce leaks or memory issues. I don't pretend that the code is free of bugs but at least these mechanisms are in place in the development cycle of the project. ## Compilation The bindings rely on [autocxx](https://google.github.io/autocxx/) to automatically generate rust FFI code from existing C++ include file. Autocxx is powerful but it can fail to process complex headers like `LIEF/ELF/Binary.hpp`. Thus, I had to create some kind of wrapper over the existing `LIEF/*.hpp` header files such as autocxx can process them. These wrappers are available in the directory `api/rust/include/` ![overview](https://lief.re/blog/2024-04-28-rust/design.webp) The time to generate the Rust FFI code for the different C++ headers is significant: about ~50s with the current bindings. This generation time can be problematic for the end user especially if LIEF is indirectly imported from other dependencies. On the other hand, for fixed versions of LIEF, cxxgen and, autocxx, the code generated by cxxgen and autocxx is *always* the same. Thus, we can pregenerate and precompile these files to save time during the pure-rust compilation step. ## Docker All the different steps mentioned in the previous parts: pre-compilation, ASAN, and code coverage are CI-compiled and **fully Dockerized**. It might also be worth mentioning that the pre-compiled FFI artifacts are also compiled and **cross-compiled** with Docker. Yes, cross-compiled. ### Cross-Compilation & CI **Digression** Feel free to skip this part which is not strictly related to LIEF & Rust. LIEF uses [GitHub Actions](https://github.com/lief-project/LIEF/tree/main/.github/workflows) for CI. From my experience, macOS and Windows runners are *less* available than Linux runners (i.e. you wait more for these runners). In addition, if you use these runners for a private repository (which is not the case for LIEF), you have a pool of 2000 minutes for the CI of the private repo. Depending on the runner you are using, these minutes are counted with a multiplier[^gh-multiplier]: | Operating system | Minute multiplier | |------------------|-------------------| | Linux | 1 | | Windows | 2 | | OSX | 10 | **1 minute** spent on a **macOS runner** is equivalent to **10 minutes** spent on a Linux runner. Hence, if your private project is exclusively using the macOS runner, you don't have 2000 minutes (~33h) but 200 minutes (~3h). And then, after this pool of 2000 minutes, 1 minute on a macOS 6 vCPU is priced at **`0.16$`** while the same minute on a Linux 8 vCPU is priced at **`0.032$`**. Given those facts, cross-compiling for macOS and Windows can be interesting. LLVM provides all the facilities to perform this cross-compilation[^ad-hoc] and since we are only generating static libraries, we don't even need the libraries for these platforms. **So yes, LIEF core and the Rust FFI library are cross-compiled for Windows(MT/MD CRT) and OSX(aarch64, x86_64) with a Docker container running on Linux :)** The Windows and OSX runners are only used for testing that the cross-compilation worked well (i.e. `ld64` can `link.exe` can link the cross-compiled libraries) and that the test suite is also working. Long story short, we save resources and CI minutes by cross-compiling for Windows and OSX in a Docker running on a Linux runner. As a side effect, we also get fully reproducible builds. The whole pipeline (LIEF core compilation, ASAN, coverage, S3 upload) takes less than 15 minutes (with cache optimizations). ## Other Projects LIEF Rust bindings might not be suitable for all the projects. Especially, if you are looking for a pure-safety-rust library or a `#![no_std]` context, please consider using these alternatives which are the standards libraries in Rust: * Goblin: https://github.com/m4b/goblin * gimli-rs - object: https://github.com/gimli-rs/object ## Acknowledgment Thank you to Erynian for the initial introduction of autocxx when I was working at Quarkslab :wink: [^gh-multiplier]: https://docs.github.com/en/billing/managing-billing-for-github-actions/about-billing-for-github-actions#minute-multipliers [^ad-hoc]: Including the generation of an ad-hoc signature for the Apple Silicon binaries (c.f [ld/MachO/SyntheticSections.cpp](https://github.com/llvm/llvm-project/blob/llvmorg-18.1.0-rc4/lld/MachO/SyntheticSections.cpp)) --- # LIEF v0.14.0 > LIEF 0.14.0 release highlights: faster Python bindings, updates to ELF and PE, and ongoing work on Rust, DWARF, and PDB support. - Canonical URL: https://lief.re/blog/2024-01-20-lief-0-14-0/ - Markdown: https://lief.re/blog/2024-01-20-lief-0-14-0/index.md - Authors: Romain Thomas - Published: 2024-01-21T00:00:00Z - Modified: 2024-01-21T00:00:00Z - Tags: release, Python, ELF, PE, Rust, DWARF, PDB LIEF v0.14.0 is out, here is an overview of the main changes! ## What's new? ### Python Bindings LIEF v0.14.0 comes with some internal enhancements for the bindings. First, LIEF now uses [nanobind](https://nanobind.readthedocs.io/en/latest/) instead of Pybind11. This change is motivated by the fact that `nanobind` reduces the compilation time while also improving the overall performances of the bindings[^nanobind-why]. The typing stubs (`.pyi`) are almost complete. This means that almost all the functions and classes have accurate typing information that is not `object` or `Any`. Finally, `setuptools` has been replaced by [scikit-build-core](https://github.com/scikit-build/scikit-build-core) as it provides a cleaner API to generate native wheels. ### ELF LIEF's ELF module now supports the GNU properties notes and exposes a friendly API to access the underlying properties information. For instance, one can check if AArch64's PAC is used by an ELF binary using the following API: ```python import lief elf = lief.ELF.parse("aarch64-binary.elf") prop: lief.ELF.NoteGnuProperty = elf.get(lief.ELF.Note.TYPE.GNU_PROPERTY_TYPE_0) aarch64_feat: lief.ELF.AArch64Feature = prop.find(lief.ELF.NoteGnuProperty.Property.TYPE.AARCH64_FEATURES) if lief.ELF.AArch64Feature.FEATURE.PAC in aarch64_feat.features: print("PAC is supported!") ``` In addition, the ELF parser can be tweaked to disable parsing some specific parts of an ELF file. For instance, one can skip parsing the relocations as follows: ```python import lief config = lief.ELF.ParserConfig() config.parse_relocations = False # ELF object without relocations information elf = lief.ELF.parse("some-binary.elf", config) ``` ### PE As of now, one of the major design issues in LIEF is the enum API. Indeed, when I started to develop LIEF, I wanted to have class and enum names as close to their names mentioned in official documentation. But it turned out that those names are -- sometimes -- already `#define` in system headers. It means that including a system header which already defines one of these names causes a compilation error: ```cpp #include // #define IMAGE_FILE_MACHINE_AM33 0x01d3 #include // /!\ Compilation error on IMAGE_FILE_MACHINE_AM33 ``` The current (hacky) workaround for this issue is a `undef.h` file which `#undef` the names that create conflict between system definition and LIEF (c.f. `LIEF/PE/undef.h`). Yes, it's a hack and the current ongoing work to address this issue is a complete refactoring of the enums API which starts with a re-scoping. Currently, **all** the enums are defined in a **single** header file and some of them are used by only one class. For instance, the enum [`LIEF::PE::SIG_ATTRIBUTE_TYPES`](https://github.com/lief-project/LIEF/blob/2d9855fc7f9d4ce6325245f8b75c98eb7663db60/include/LIEF/PE/enums.hpp#L1315-L1330), has been re-scoped in the `LIEF::PE::Attribute`: ```cpp // Before (v0.13.x): LIEF/PE/enums.hpp enum class SIG_ATTRIBUTE_TYPES { UNKNOWN = 0, CONTENT_TYPE, ... }; // Now (v0.14.0): LIEF/PE/signature/Attribute.hpp class LIEF_API Attribute : public Object { public: enum class TYPE { UNKNOWN = 0, CONTENT_TYPE, ... }; } ``` As of LIEF `v0.14.0`, the PE format is mostly impacted by this refactoring and the other formats should be progressively updated accordingly. ## On Going Work ***As a reminder, LIEF is exclusively developed on my spare time, so some functionalities might take time to be completed and integrated*** ### Rust Bindings This is still ongoing and the bindings are almost completed for ELF, PE, and Mach-O. I still need to create the bindings for the enums and figure out a way to reduce the compilation time but it keeps moving! ### DWARF & PDB LIEF will welcome DWARF and PDB debug information support through an **external** extension. This module will provide a comprehensive API to iterate over DWARF & PDB information. ## Final Word Since LIEF 0.13.2, this new version introduces **274** new commits, with **35 292** additions and **39 392** deletions thanks to **15** contributors! The [complete changelog](https://lief-project.github.io/doc/stable/changelog.html#january-20-2024) is also available. Thank you also to F., [antipatico](https://github.com/antipatico), and [MobSF](https://github.com/MobSF) for their sponsoring. [^nanobind-why]: https://nanobind.readthedocs.io/en/latest/why.html --- # LIEF v0.13.0 > LIEF 0.13.0: in-memory Mach-O parsing, framed ELF sections, an exception-free and RTTI-free core, PEP 621 Python packaging, and generated .pyi type stubs. - Canonical URL: https://lief.re/blog/2023-04-09-lief-0-13-0/ - Markdown: https://lief.re/blog/2023-04-09-lief-0-13-0/index.md - Authors: Romain Thomas - Published: 2023-04-09T00:00:00Z - Modified: 2023-04-09T00:00:00Z - Tags: release, Mach-O, ELF, Python LIEF v0.13.0 is eventually out! The full changelog is available [here](https://lief-project.github.io/doc/stable/changelog.html#april-9-2023), but here is a summary of the main changes. ## Mach-O in Memory Parser LIEF is now able to parse a Mach-O file from an in-memory pointer: ```cpp uintptr_t mhdr = ...; // Absolute address to an in-memory Mach-O file auto macho = LIEF::MachO::Parser::parse_from_memory(mhdr); ``` This feature can be handy on iOS to access the in-memory content of the binary. Firstly because some parts of the application code are encrypted (thanks to the ``LC_ENCRYPTION_INFO`` commands). Secondly, it can be used on protected code that fills the ``__data`` segment with clear (original) strings (c.f. [Gotta Catch 'Em All: Frida & jailbreak detection](https://www.romainthomas.fr/post/21-07-pokemongo-anti-frida-jailbreak-bypass/#what-about-lief)). This feature could also be used in pair with ``_dyld_get_image_header``, once we are injected into the targeted process: ```cpp size_t count = _dyld_image_count(); for (size_t i = 0; i < count; ++i) { llvm::StringRef Name = _dyld_get_image_name(i); if (Name.contains("MyApp.app/Target")) { auto* mhdr = _dyld_get_image_header(i); auto macho = MachO::Parser::parse_from_memory((uintptr_t)mhdr); } } ``` It might also work to access the dyld shared cache libraries but this aspect does not have been heavily tested. ## Framed ELF Sections The ELF format is -- by far -- the most tricky format especially when it comes dealing with sections and segments. These ELF structures are two different ways of *slicing* the binary data: 1. Sections are used by the compiler and the linker 2. Segments are used by the system loader. LIEF implements a mechanism to deal with this dual representation so that if the user updates the ``.text`` section, the changes are also committed in the associated segment (if any). It turns out that in some scenarios[^PST], we might want to NOT *commit* the changes we are doing on the sections. The ``lief.ELF.Section.as_frame()`` function can be used to make the section "*frame only*". All the attributes of the section will be committed in the final ELF binary but LIEF won't consider this section to write the content in the binary. One can use -- for instance -- this function to corrupt sections attributes: ```python elf = lief.parse("/bin/ls") text = elf.get_section(".text").as_frame() text.offset = 0xffffff elf.write("ls.modified") ``` As the `.text` section is set as "framed", its `0xffffff` offset is not considered for changing its content. Thus, this code does not update anything: ```python text.content = ... ``` Since the ELF loader only relies on segment, this modification does not affect the execution of the modified binary. ## Internal Changes As announced in the previous v0.12.0 release, the LIEF codebase is now free from exceptions and RTTI. This means that the core library can be compiled in an `-fno-exceptions` context. The Python build process is now compliant with the [PEP 621](https://peps.python.org/pep-0621/) ``pyproject.toml`` requirement. You can find this file, along with `config-default.toml`, in the ``api/python`` directory. The `config-default.toml` file can be used to tweak the compilation of the bindings. This is essentially an interface over the LIEF CMake option. Lastly, the `setup.py` file used for compiling the binding has moved from the root directory to `api/python`. On top of that, the Python bindings also generate stub interfaces (`.pyi`) which are handy for type checking and code completion (cf. [issues/650](https://github.com/lief-project/LIEF/issues/650)). ## Final Words Some features like the Rust bindings are not released yet they still require some ongoing work. I also really do hope to be able to work on enhancing the PE format modification in the next release but this will be balanced with the spare time I have to work on it :) Thank you for using LIEF! [^in-mem]: https://www.romainthomas.fr/post/21-07-pokemongo-anti-frida-jailbreak-bypass/#what-about-lief [^PST]: https://passthesalt.ubicast.tv/videos/the-poor-mans-obfuscator/ --- # Mach-O Support Enhancements > Mach-O rewriting in LIEF: __LINKEDIT fixes, chained fixups and exports trie support, converting a binary into a dylib, adding exported symbols, and code injection. - Canonical URL: https://lief.re/blog/2022-05-08-macho/ - Markdown: https://lief.re/blog/2022-05-08-macho/index.md - Authors: Romain Thomas - Published: 2022-05-08T00:00:00Z - Modified: 2022-05-08T00:00:00Z - Tags: Mach-O, macOS, iOS, code-injection **tl;dr** The next release of LIEF (v0.13.0) is fixing several Mach-O layout issues when adding new sections/segments. I also added the support for the two new load commands: 1. `LC_DYLD_CHAINED_FIXUPS` 2. `LC_DYLD_EXPORTS_TRIE` The support of LIEF for modifying Mach-O binaries was mostly limited to adding new load commands and thus, extending the load commands table. The [tutorial #11](https://lief.re/doc/latest/tutorials/11_macho_modification.html) explains the technical details to extend the load commands table which consists in shifting the content right after the load commands table and patching the relocations accordingly. Nevertheless, the Mach-O binaries generated by LIEF after the modifications were somehow inconsistent regarding ``codesign``. As a consequence, the binaries generated by LIEF could not be signed and executed on iOS or -- more recently -- an Apple M1. ![Technical diagram](https://lief.re/img/stockholm/Layout/Layout-4-blocks.svg) ** LIEF is now able to generate Mach-O-modified files that can be signed and that follow a strict layout, enforced by dyld and codesign. ** To better understand what was wrong, let's consider the following script in which we add two new segments: ```python import lief target = lief.parse("mbedtls_selftest_arm64.bin") segment = lief.MachO.SegmentCommand("__NEW", [0] * 0x123) target.add(segment) segment = lief.MachO.SegmentCommand("__NEW", [0] * 0x456) target.add(segment) target.write("test.out") ``` Under the hood, LIEF was relocating the binary to add two new `LC_SEGMENT` commands and was allocating space **at the end** of the file to store the content of the new segments. In particular, the new segments data were located **after** the content of the `__LINKEDIT` segment which breaks the layout required by `codesign`. The following figure depicts the layout of a Mach-O file from the original layout to the layout generated by LIEF v0.13.0. ![Technical diagram](https://lief.re/blog/2022-05-08-macho/macho_layout.svg) In LIEF v0.13.0 we fixed this inconsistency to make sure that the content of the new segments are located **before** the content of the `__LINKEDIT` segment. We can perform this change without breaking the binary as `__LINKEDIT` is a kind of self-contained *blob of data*[^instrplace]. `codesign` requires the `__LINKEDIT` segment at the end of the file because the signature is appended at the end of the file. Otherwise, `codesign` would have to perform the similar *relocation* process done by LIEF. ## ``__LINKEDIT`` The ``__LINKEDIT`` segment plays an important role in the layout of the Mach-O format and its execution. This segment is used to store information about the exports, the symbols, the relocations, the signature, and more broadly, information used by the `dyld` loader to load the binary. This segment has a known layout which is described in the following figure: ![Technical diagram](https://lief.re/blog/2022-05-08-macho/linkedit_layout.svg) This layout is very strict and its content must follow the same order as mentioned in the previous figure. In addition, there are sanity checks that ensure all the `__LINKEDIT`'s chunks are contiguous within the `__LINKEDIT` content. If the layout is wrong, the executable **could** run but it won't likely pass the `codesign` checks. This strict layout can be seen -- at first sight -- as a major hurdle for modifying Mach-O files but since the `__LINKEDIT` segment is located **at the end** of the file, we can extend it or shrink it quite easily. ![Technical diagram](https://lief.re/img/stockholm/Layout/Layout-4-blocks.svg) ** LIEF v0.13.0 is able to regenerate the content of this segment from the LIEF objects stored in the LIEF::MachO::Binary object ** Completely regenerating the __LINKEDIT segment enables to perform advanced modifications like creating exports and adding or removing symbols as it is discussed in the next sections. ## `LC_DYLD_CHAINED_FIXUPS & LC_DYLD_EXPORTS_TRIE` Compared to the ELF and PE formats, the relocations and the exported functions of Mach-O binaries are not wrapped by a *table of entries* In the Mach-O format, the relocations are encoded either: 1. By a *bytecode* located in the `LC_DYLD_INFO` command 2. By a *chained fixups* located in the `LC_DYLD_CHAINED_FIXUPS` On the other hand, the exports are encoded in a [Trie](https://en.wikipedia.org/wiki/Trie) located either 1. In the `LC_DYLD_INFO` command 2. In the `LC_DYLD_EXPORTS_TRIE` `LC_DYLD_CHAINED_FIXUPS` appeared more recently compared to the `LC_DYLD_INFO` command for which the differences are described in the blog post: [*How iOS 15 makes your app launch faster*](https://www.emergetools.com/blog/posts/iOS15LaunchTime). The `LC_DYLD_EXPORTS_TRIE` has the same structure as `LC_DYLD_INFO[Export Trie]` but the export information has been moved in this dedicated load command. ## Converting a Mach-O Binary into a Library Converting a binary into a library can be useful to harness a fuzzed binary or to instrument/debug a specific function in a controlled environment (like an unknown cryptography function or a whiteboxed function) In the [tutorial #8](https://lief-project.github.io/doc/latest/tutorials/08_elf_bin2lib.html), we described the process to perform this transformation on an ELF binary and the transformation for a Mach-O binary is a bit more straightforward. Let's consider the following code: ```c #include #include #include static int X = 1; int compute() { return X++; } int main(int argc, const char** argv) { for (size_t i = 0; i < argc; ++i) { printf("compute(): %d\n", compute()); } return 0; } ``` It can be compiled with: ```bash romain@Mac-M1 % clang -O3 -fvisibility=hidden -Wl,-x -o bin2lib.bin bin2lib.c ``` Which produces this executable: [bin2lib.bin](https://lief.re/blog/2022-05-08-macho/bin2lib.bin) To convert this binary into a library, we first need to change its type in the Mach-O's header: ```python import lief bin2lib = lief.parse("bin2lib.bin") bin2lib.header.file_type = lief.MachO.FILE_TYPES.DYLIB bin2lib.write("bin2lib.dyld") ``` It's should be technically enough, but `dyld_info` raises some concerns: ```bash romain@Mac-M1 % dyld_info ./bin2lib.dylib dyld_info: './bin2lib.dylib' in './bin2lib.dylib' MH_DYLIB is missing LC_ID_DYLIB ``` This can be confirmed by looking at the source code of [dyld](https://github.com/apple-oss-distributions/dyld/blob/5c9192436bb195e7a8fe61f22a229ee3d30d8222/common/MachOAnalyzer.cpp#L775-L779). To fix this error, we just have to create a new `LC_ID_DYLIB` command: ```diff {style=pastie} import lief bin2lib = lief.parse("bin2lib.bin") bin2lib.header.file_type = lief.MachO.FILE_TYPES.DYLIB + bin2lib.add(lief.MachO.DylibCommand.id_dylib("bin2lib.dylib", 0, 1, 2)) bin2lib.write("bin2lib.dyld") ``` Which enables to dlopen `bin2lib.dyld` ```python import ctypes handler = ctypes.cdll.LoadLibrary("bin2lib.dyld") # ``` ## Adding Symbols Thanks to the improvements on the `__LINKEDIT` segment, we can now create new exports. If we consider the stripped function `int compute()` from the binary in the previous section, we can create a new export as follows: ![Technical diagram](https://lief.re/blog/2022-05-08-macho/function_to_export.svg) ```python address = 0x100003f18 original.add_exported_function(address, "_compute") ``` ## Code Injection Another use case of these improvements is the capability to inject code in Mach-O file **and to re-sign** the modified binary. Code signing is not required for x86-64 binaries but it becomes mandatory when targeting the arm64 architecture. Let's consider the library `_heapq.cpython-39-darwin.so` which is one of the first libraries dynamically loaded by the Python interpreter. The injection consists in: 1. Creating new segments in the library `_heapq.cpython-39-darwin.so` that will embed our shellcode 2. Changing the address of one of the exported functions to redirect the execution to the shellcode's entrypoint. By running the python interpreter with the environment variable ``DYLD_PRINT_APIS=1`` we can observe the following output: ```bash romain@Mac-M1 ~ % DYLD_PRINT_APIS=1 python3 -c "import io" dyld[76439]: _dyld_is_memory_immutable(0x1b3f8cea0, 26) => 1 dyld[76439]: dlopen("/opt/homebrew/Cellar/python@3.9/3.9.5/Frameworks/Python.framework/Versions/3.9/lib/python3.9/lib-dynload/_heapq.cpython-39-darwin.so", 0x00000002) dyld[76439]: dlopen(_heapq.cpython-39-darwin.so) => 0x208f35800 dyld[76439]: dlsym(0x208f35800, "PyInit__heapq") dyld[76439]: dlsym("PyInit__heapq") => 0x104bcb824 ``` It suggests that ``PyInit__heapq`` is a suitable function for redirecting the execution to the shellcode's entrypoint. To create the shellcode, we can use [gdelugre/shell-factory](https://github.com/gdelugre/shell-factory) developed by a former colleague and which provides **no less than a C++ STL-like** to create shellcode. Thanks to this project, we can create the following shellcode: ```cpp volatile uintptr_t ORIGINAL_EP = 0xdeadc0de; volatile uintptr_t IMAGEBASE = 0x00c0de; using PyInit__heapq_t = void(*)(); inline uintptr_t imagebase() { /* * The value of IMAGEBASE is set by the injector. * After the patch, it contains the relative virtual address of &IMAGEBASE * in the final binary. */ return reinterpret_cast(&IMAGEBASE) - IMAGEBASE; } SHELLCODE_ENTRY { uintptr_t base = imagebase(); Pico::printf("LIEF says hello!\n"); Pico::printf("Time to jump on the real function: %p\n", ORIGINAL_EP); auto PyInit__heapq = reinterpret_cast(base + ORIGINAL_EP); return PyInit__heapq(); } ``` **`Pico::printf`** The attentive reader may have noticed the `Pico::printf("[...] %p")` which is correctly supported by shell-factory (see: [include/pico/format.h](https://github.com/gdelugre/shell-factory/blob/25639dd517ace9a9292db38f8ca423808317de65/include/pico/format.h)) The compiled shellcode can be downloaded here: [lief_demo_darwin_arm64.bin](https://lief.re/blog/2022-05-08-macho/lief_demo_darwin_arm64.bin). To inject the shellcode in `_heapq.cpython-39-darwin.so`, we first need to copy the shellcode's segments in the library: ```python shellcode = lief.parse("lief_demo_darwin_arm64.bin") heapq = lief.parse("_heapq.cpython-39-darwin.so") for segment in shellcode.segments: seg_name = segment.name.replace("__", "") seg = lief.MachO.SegmentCommand(f"__L{new_seg_name}", list(segment.content)) heapq.add(new_seg) ``` Then, we have to patch the Mach-O exports trie to change the address of `PyInit__heapq` to the shellcode's entrypoint: ```python shellcode_rva_entry = ... for exp in heapq.dyld_info.exports: if exp.symbol.name != "_PyInit__heapq": continue original = exp.address exp.address = shellcode_rva_entry return original ``` Finally, we can rewrite the library: ```python heapq.write("_heapq.cpython-39-darwin.so.patched") ``` and **sign it**: ```bash romain@Mac-M1 ~ % codesign -f --verbose -s - _heapq.cpython-39-darwin.so.patched ``` Now when running the Python interpreter, we can observe the execution of the shellcode: ```bash romain@Mac-M1 ~ % python3 LIEF says hello! Time to jump on the real function: 0x15f8 Python 3.9.5 (default, May 3 2021, 19:12:05) [Clang 12.0.5 (clang-1205.0.22.9)] on darwin Type "help", "copyright", "credits" or "license" for more information. >>> ``` **Injection** The script that contains the complete logic of the transformation is available [here](https://gist.github.com/romainthomas/16f384a21fe408c7d20e369d75e69588) and, `_heapq.cpython-39-darwin.so.patched` can be downloaded [here](https://lief.re/blog/2022-05-08-macho/_heapq.cpython-39-darwin.so.patched). Surprisingly, we open the **patched** version of the library ([`_heapq.cpython-39-darwin.so.patched`](https://lief.re/blog/2022-05-08-macho/_heapq.cpython-39-darwin.so.patched)) in IDA and we jump on the symbol `_PyInit__heapq`, it actually displays this function: [*IDA Version 7.7.211224*, January 18, 2022](https://hex-rays.com/products/ida/news/) ![IDA Output when jumping on _PyInit__heapq](https://lief.re/blog/2022-05-08-macho/IDA_patched_pyinit_heapq.png) Which is the original function and not the function associated with the shellcode whilst the patched library prints `LIEF says hello [...]` On the other hand, if we get the address of `_PyInit__heapq` with LIEF: ```python import lief patched = lief.parse("./_heapq.cpython-39-darwin.so.patched") symbol = patched.get_symbol("_PyInit__heapq") print(hex(symbol.export_info.address)) ``` The result is: ![Technical diagram](https://lief.re/img/stockholm/General/Thunder-move.svg) **`_PyInit__heapq: 0xf824`** Jumping on this address gives a better output (once manually disassembled): ![IDA Output when jumping on 0xf824](https://lief.re/blog/2022-05-08-macho/IDA_patched_pyinit_heapq_code.png) We recognize the shellcode's entrypoint function . What's happened in IDA since this is the function located at `0xf824` which is executed and thus, resolved by `dyld` and not IDA? IDA is confused because Mach-O's symbols can be stored in two different commands: 1. `LC_DYLD_INFO.export_trie` / `LC_DYLD_EXPORTS_TRIE` 2. `LC_SYMTAB` `LC_DYLD_INFO.export_trie` / `LC_DYLD_EXPORTS_TRIE` are used to store the **exported symbols** while `LC_SYMTAB` stores symbols for other purposes. The important point is that the same symbol can be duplicated in these two commands **with different addresses**. ![Technical diagram](https://lief.re/img/stockholm/General/Thunder-move.svg) ** IDA gives the priority to the `LC_SYMTAB` over the exports trie while the Mach-O loader uses the exports trie. ** The following figure illustrates why it can be confusing: ![Technical diagram](https://lief.re/blog/2022-05-08-macho/patched.svg) Actually, I intentionally took a shortcut in the LIEF script that resolves the address of `_PyInit__heapq` and we can programmatically access these two addresses as follows: ```diff {style=pastie} import lief patched = lief.parse("./_heapq.cpython-39-darwin.so.patched") symbol = patched.get_symbol("_PyInit__heapq") + print(hex(symbol.value)) print(hex(symbol.export_info.address)) + # 0x15f8 address from the LC_SYMTAB # 0xf824 address from the export trie ``` ![Technical diagram](https://lief.re/img/stockholm/General/Attachment-2.svg) ** We can observe a similar issue with BinaryNinja, Ghidra and, to a lesser extent, Radare2 ** ## BinaryNinja *Version 3.0* ![BinaryNinja Result](https://lief.re/blog/2022-05-08-macho/binaryninja_result.png) ## Ghidra [*Version 10.1.2*](https://github.com/NationalSecurityAgency/ghidra/releases/tag/Ghidra_10.1.2_build) - Jan 26, 2022 ![Ghidra Result](https://lief.re/blog/2022-05-08-macho/ghidra_result.png) ## Radare2 [*Version: 5.6.6*](https://github.com/radareorg/radare2/releases/tag/5.6.6) - Mar 22, 2022 ```bash $ r2 _heapq.cpython-39-darwin.so.patched [0x00000000]> aaa ... [0x00000000]> ia [Imports] nth vaddr bind type lib name ――――――――――――――――――――――――――――――――― 0 0x000021ec NONE FUNC PyErr_SetString 1 0x00000000 NONE FUNC PyExc_IndexError 2 0x00000000 NONE FUNC PyExc_RuntimeError 3 0x00000000 NONE FUNC PyExc_TypeError 4 0x000021f8 NONE FUNC PyList_Append 5 0x00002204 NONE FUNC PyList_SetSlice 6 0x00002210 NONE FUNC PyModuleDef_Init 7 0x0000221c NONE FUNC PyModule_AddObject 8 0x00002228 NONE FUNC PyObject_RichCompareBool 9 0x00002234 NONE FUNC PyUnicode_FromString 10 0x00002240 NONE FUNC _PyArg_CheckPositional 11 0x0000224c NONE FUNC _Py_Dealloc 12 0x00000000 NONE FUNC _Py_NoneStruct 13 0x00000000 NONE FUNC dyld_stub_binder [Exports] nth paddr vaddr bind type size lib name ――――――――――――――――――――――――――――――――――――――――――――――――――― 0 0x000015f8 0x000015f8 GLOBAL FUNC 0 _PyInit__heapq ``` On the other hand, the `afl` command outputs a better result: ```bash [0x00000000]> afl 0x000015f8 1 12 sym._PyInit__heapq 0x00001604 6 108 sym._heapq_exec 0x00002238 1 8 fcn.00002238 0x00002220 1 8 fcn.00002220 0x00002250 1 8 fcn.00002250 ... 0x0000f824 1 88 sym.imp._PyInit__heapq ``` ### Demo [Terminal recording](https://lief.re/blog/2022-05-08-macho/out.rec) ## Conclusion These changes strengthen LIEF to read and modify Mach-O binaries. It should enable to develop and create new reverse engineering and binary analysis techniques. For those who are interested in Mach-O (and ELF) *tricks* that could prevent static analysis tools from working correctly, I'll present *The Poor Man's Obfuscator* at [Pass The Salt](https://2022.pass-the-salt.org/) in July 2022 :) [^instrplace]: In the general case, we can't insert content between two arbitrary segments as it could break the binary. For instance, if the `__TEXT` segment references variables in the `__DATA` segment with relative addressing, inserting some data between these two segments will likely break the relative addressing. --- # LIEF v0.12.0 > LIEF 0.12.0: PE rich header and checksum recomputation, span-based section content with memoryview in Python, and the start of the exception-free refactoring. - Canonical URL: https://lief.re/blog/2022-03-27-lief-v0-12-0/ - Markdown: https://lief.re/blog/2022-03-27-lief-v0-12-0/index.md - Authors: Romain Thomas - Published: 2022-03-27T00:00:00Z - Modified: 2022-03-27T00:00:00Z - Tags: release, PE, Python, C++ We are thrilled to announce that LIEF v0.12.0 is released! You can find the complete changelog [here](https://lief-project.github.io/doc/stable/changelog.html#march-25-2022). ## LIEF v0.12.0: What's New? LIEF v0.12.0 is a balanced mix of new features, internal refactoring, and performance improvement. ### New Features Regarding the new features, we added support for recomputing the PE's rich header and the PE's checksum. The PE's rich header is a well-known-hidden[^rheader] feature that can be helpful to fingerprint a PE binary. LIEF enables -- since the version `v0.7.0` -- to access this part of the PE file with the following API: ```python import lief pe_file = lief.parse("hello.exe") rich_header = pe_file.rich_header print(f"XOR Key: {rich_header.key}") for e in rich_header.entries: print(f"{e.id}: {e.build_id} {e.count}") ``` In LIEF `v0.12.0`, we added two functions: 1. `LIEF::PE::RichHeader::raw`: To generate the rich header blob with or without an XOR key. 2. `LIEF::PE::RichHeader::hash`: To generate the MD5/SHA-1/SHA-256/(...) of the rich header blob. For those who are looking for PE's markers or tracking PE binaries, these two functions could be used to generate a characteristic of the binary, regardless of the xor-key: ```python # [...] rich_header = pe_file.rich_header marker = bytes(rich_header.hash(lief.PE.ALGORITHMS.SHA_1)).hex() ``` Still about the PE format, we added `LIEF::PE::OptionalHeader::computed_checksum()` which returns the recomputed value of the PE's checksum (`LIEF::PE::OptionalHeader::checksum()`). For regular binaries, the verification of the OptionalHeader's checksum is not enforced by Windows and the integrity checks are usually deferred to the PE's Authenticode. Nonetheless, verifying the `checksum()` value with the output of `computed_checksum()` could help identify binaries that would have been modified after the compilation. Finally, we added the support for the PE's delayed imports in LIEF and Luca Moro added the support of the `LC_FILESET_ENTRY` command in the Mach-O format. ### Refactoring & Performance Improvement We also refactored and enhanced LIEF's internal codebase. Among those changes, we started to get rid of the C++ exceptions as described in this blog post: [LIEF RTTI & Exceptions](https://lief.re/blog/2022-02-13-lief-rtti-exceptions/) We also introduced a ``std::span`` like interface (based on [tcbrindle/span](https://github.com/tcbrindle/span)) to avoid returning and potentially copying ``std::vector``. For instance, [LIEF::Section::content](https://github.com/lief-project/LIEF/blob/57294452a1470f2e1432112d1649068a8cd8047e/include/LIEF/Abstract/Section.hpp#L50) now uses the span interface. In the Python API, functions and properties that bind a span-returning function now return a ``py::memoryview`` instead of the *list of bytes*. The original *list of bytes* can be recovered as follows: ```python bin = lief.parse("/bin/ls") section = bin.get_section(".text") if section is not None: memory_view = section.content list_of_bytes = list(memory_view) ``` About the performances, we did a global refactoring of the ELF builder as described in this blog post: [New ELF Builder](https://lief-project.github.io/blog/2022-01-23-new-elf-builder). We also reduced the memory footprint of the ELF parser. For instance, in LIEF v0.11.5 a binary of 1.5G takes 3G or RAM[^oups] while in LIEF v0.12.0, it takes quite the same memory as the file size. ![Memory profiling](https://lief.re/blog/2022-03-27-lief-v0-12-0/memory_profile.png) [Eric Kilmer](https://github.com/ekilmer) also did a nice and complete cleaning of the [LIEF CMake integration](https://github.com/lief-project/LIEF/pull/674) In February 2022, [tmp.0ut v2](https://tmpout.sh/2/) has been released and [@netspooky](https://twitter.com/netspooky) presented interesting tricks on the ELF format [^elf_parser] [^elf_endianness]. We fixed the ELF parser to make sure we handle these tricks. ## What's Next? We started to implement Rust bindings for LIEF thanks to [cxx](https://cxx.rs/) and [google/autocxx](https://github.com/google/autocxx). These bindings are in their early stages and we can't confirm they will be present in the next release. In the current development stage, the API looks like this: ```rust let mut path: String = "/bin/ls"; match Binary::parse(&path) { Binary::ELF(elf) => { println!("ELF binary"); for segment in elf.segments() { println!("Address: {:x}", segment.virtual_address); } }, Binary::PE(pe) => { println!("PE binary"); let text_section = pe.get_section(".text"); text_section.name = ".foo"; text_section.file_offset = 0x123; text_section.commit(); // Commit the changes }, Binary::MachO(macho) => { println!("MachO binary"); for command in macho.commands() { match command { Commands::Dylib(dylib) => { ... }, Commands::Main(main_cmd) => { ... }, } } }, Binary::Unknown(x) => { println!("Unknown"); }, } ``` We will also merge the (still private) branch that enables to parse Mach-O from memory as well as the global improvement of the Mach-O's builder. Regarding LIEF's experimentations and work in progress, here is a list of topics on which we are working or we would like to work: | Topic | Status | |---------------------------------------------------------------------------------|:------------------------------------| | Parsing ELF files from memory | Not started yet | | Parsing DART/Flutter snapshots | PoC | | Creating an ELF from scratch | PoC | | Parsing PE's private Authenticode: MS Counter Signature | Not started yet | | Refactoring the PE's builder | Not started yet, priority undefined | | Supporting the Mach-O's commands: LC_DYLD_CHAINED_FIXUPS / LC_DYLD_EXPORTS_TRIE | Done, under testing | | Supporting the archive format (AR) | Early stage | If you are interested in supporting some of these topics, feel free to reach out. Enjoy! [^rheader]: https://www.virusbulletin.com/virusbulletin/2020/01/vb2019-paper-rich-headers-leveraging-mysterious-artifact-pe-format/ [^elf_parser]: https://tmpout.sh/2/3.html [^elf_endianness]: https://tmpout.sh/2/14.html [^oups]: More generally, we have a factor 2 in memory compared to the file size. --- # LIEF RTTI & Exceptions > Why LIEF is removing C++ exceptions and RTTI: error-handling costs, the has_/get_ pattern, and the move to result-based APIs and LLVM-style RTTI. - Canonical URL: https://lief.re/blog/2022-02-13-lief-rtti-exceptions/ - Markdown: https://lief.re/blog/2022-02-13-lief-rtti-exceptions/index.md - Authors: Romain Thomas - Published: 2022-02-13T00:00:00Z - Modified: 2022-02-13T00:00:00Z - Tags: C++, internals, error-handling ## try { {#try} When we started to develop LIEF, we choose to manage errors through the C++ exceptions as it is widely spread in Java. However, with a little hindsight it was not the best choice in the design of LIEF. First off, LIEF is a **library** and the API functions that throw exceptions are not compatible with library's users that are not using exceptions (e.g with the ``-fno-exceptions`` flag). It is also considered as a bad practice quoting from C++ Coding Standards: **C++ Coding Standards: Item 62** *“Don't throw stones into your neighbor's garden: There is no ubiquitous binary standard for C++ exception handling.”* For instance, the function ``LIEF::ELF::Binary::get_section(const std::string& name)`` threw an exception if the section were not found. To avoid raising the exception, the API exposes helpers that can be used to check -- beforehand -- that it will not take the *exception path*: ```cpp if (bin.has_section(".toto")) { auto& sec = bin.get_section(".toto"); // Ok no exception } // With exception: try { bin.get_section(".toto"); } catch (const std::exception&) { // .toto does not exist :( } ``` Actually, the `has_` / `get_` pattern hides another issue: **the performances**. Basically, `has_section(...)` iterates over the list of the sections to check if a section with the given name exists and ``get_section()`` iterates **again** on this list to access the section. **The code performs twice the same iteration**. This is not a big deal for the ELF sections as they are quite small but it can be problematic for large sequences like the symbols table. In LIEF v0.12.0 we changed the API of these functions to return **a pointer** on these objects instead of **a reference**. If the item can't be found, it returns a `nullptr`. **The API contract of these functions is changing from raising an exception into returning a `nullptr`. The documentation has been updated accordingly and the list of the functions which have changed are listed [here](https://gist.github.com/romainthomas/37da45b043c5f8b8db6be2767611f625)** As a result, the previous code can be re-written as follows: ```cpp if (bin.has_section(".toto")) { auto* sec = bin.get_section(".toto"); // Non nullptr instead of a reference } // Or: if (auto* sec = bin.get_section(".toto")) { // ... } ``` This kind of API change is doable and meaningful for functions that aim at returning an **optional object** but it is less meaningful to transform a function like: ```cpp uint64_t Binary::virtual_address_to_offset(...) { ... } ``` that returns an integer (while still potentially raising an exception). In LIEF 0.12.0, this kind of function still raises an exception but in the next version (LIEF v0.13.0) the returned value will be wrapped by [**Boost's Leaf**](https://github.com/boostorg/leaf)[^leaf_note] such as the returned type will become: ```cpp // Future returned type result Binary::virtual_address_to_offset(...) { ... } // To use it: auto res = bin.virtual_address_to_offset(); if (!res) { // Error } else { uint64_t val = res.value(); // or val = *res } ``` In LIEF v0.12.0 **only internal/private functions associated with the Parser/Builder module** are using this mechanism and we plan to move to this mechanism in the public API[^cpp_11] in LIEF v0.13.0. **Warning** Boost LEAF is required in the **public headers** of LIEF. If you find conflicts, compilation issues, integration issues, or you think that it is a bad idea, **please let us know** before it becomes the default interface to manage errors. ### The RTTI LIEF also relies on the RTTI information which includes calling functions like `typeid()` or `dynamic_cast<>()`. For instance, to check if a Mach-O's LoadCommand exists, the main `MachO::Binary` class calls at some point this helper: ```cpp template bool Binary::has_command() const { static_assert(std::is_base_of::value, "Require inheritance from 'LoadCommand'"); const auto it_cmd = std::find_if( std::begin(commands_), std::end(commands_), [] (const LoadCommand* command) { return typeid(T) == typeid(*command); }); return it_cmd != std::end(commands_); } ``` This code generates extra data for the RTTI information of the LoadCommand objects which can be perfectly fine. Actually, this RTTI information is redundant as the type of a Mach-O's LoadCommand is already stored in the class itself: ```cpp class LoadCommand { ... private: LOAD_COMMAND_TYPES command_; }; ``` So instead of having this redundant RTTI information, we implemented an LLVM-like RTTI[^llvm_rtti] based on `classof()` and that uses the already present `command_` attribute. In the end, the previous `has_command()` can be updated as follows: ```cpp template bool Binary::has_command() const { static_assert(std::is_base_of::value, "Require inheritance from 'LoadCommand'"); const auto it_cmd = std::find_if( std::begin(commands_), std::end(commands_), [] (const LoadCommand* command) { return T::classof(command); }); return it_cmd != std::end(commands_); } ``` We applied this pattern for the LIEF's object where ``typeid`` was present and as a result, we managed to completely remove this function as it was redundant with an existing attribute. ## } catch (const std::length_error&) { {#catch} We welcome feedback on these changes -- whether positive or negative -- as it impacts the public API. Thank you for reading! ## } {#end} [^leaf_note]: See the section [Error Handling](https://lief-project.github.io/doc/latest/error_handling.html) of the documentation for more details. [^cpp_11]: It still keeps the public headers compliant with C++11 [^llvm_rtti]: https://llvm.org/docs/HowToSetUpLLVMStyleRTTI.html --- # New ELF Builder > After spending months on refactoring the ELF builder, here are the improvements. - Canonical URL: https://lief.re/blog/2022-01-23-new-elf-builder/ - Markdown: https://lief.re/blog/2022-01-23-new-elf-builder/index.md - Authors: Romain Thomas - Published: 2022-01-23T00:00:00Z - Modified: 2022-01-23T00:00:00Z - Tags: ELF, builder, internals, performance ## LIEF's Modification Process Let's start with a small recap of the LIEF modification process. To enable executable file formats modification, LIEF transforms the raw executable formats into an object representation. This object can be manipulated with an API that is mainly exposed through the following interfaces: | C++ | Python | |:-------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------| | [LIEF::ELF::Binary](https://lief-project.github.io/doc/latest/api/cpp/elf.html#binary) | [lief.ELF.Binary](https://lief-project.github.io/doc/latest/api/python/elf.html#binary) | | [LIEF::PE::Binary](https://lief-project.github.io/doc/latest/api/cpp/pe.html#binary) | [lief.PE.Binary](https://lief-project.github.io/doc/latest/api/python/pe.html#binary) | | [LIEF::MachO::Binary](https://lief-project.github.io/doc/latest/api/cpp/macho.html#binary) | [lief.MachO.Binary](https://lief-project.github.io/doc/latest/api/python/macho.html#binary) | Then, the LIEF's *builders* take the object representation and (*try to*) reconstruct an executable according to the user's changes. ## Challenges in Modifying ELF Binaries Compared to the PE and Mach-O formats, the ELF format is far trickier to handle for both parsing and modifying. First off, there is a strong relationship between the segment's virtual address and the file's offset associated with its content. This relationship is ruled by the following property: $$\text{\textcolor{red}{file\_offset}} \equiv \text{\textcolor{blue}{virtual\_address}} \mod{\textcolor{green}{\text{page\_size}}}$$ So basically, **we can't** insert a segment at an arbitrary virtual address. The second difficulty is about the strings table optimization that is performed on the ``.dynstr`` section. To understand how this optimization works, let's consider these two functions: ```cpp int foo() { return 1; } int call_foo() { return foo(); } ``` When these functions are compiled, the compiler generates two symbols for which the names of the symbols are referenced by the field ``st_name``. Usually, this field points in the ``.dynstr`` section: ```cpp struct Elf_Sym { Elf_Word st_name; // Offset of the symbol's name in the .dynstr section ... }; ``` Naively, we could imagine that the ``.dynstr`` section contains these two symbols names, one next to the other: ```hex 00000130 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000140 0d 00 00 00 12 00 01 00 00 00 00 00 00 00 00 00 |................| 00000150 0b 00 00 00 00 00 00 00 0a 00 00 00 12 00 01 00 |................| 00000160 0b 00 00 00 00 00 00 00 0b 00 00 00 00 00 00 00 |................| 00000170 00 74 6f 74 6f 2e 63 70 70 00 66 6f 6f 00 64 6f |.test.cpp.foo.do| 00000180 5f 66 6f 6f 00 00 00 00 10 00 00 00 00 00 00 00 |_foo............| 00000 ``` With such a layout, ``Elf_Sym("foo").st_name`` would point to the offset **0x17A** while ``Elf_Sym("do_foo").st_name`` would point to the offset **0x17E**. But the real layout of the ``.dynstr`` is a bit smaller: ```hex 00000130 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000140 0d 00 00 00 12 00 01 00 00 00 00 00 00 00 00 00 |................| 00000150 0b 00 00 00 00 00 00 00 0a 00 00 00 12 00 01 00 |................| 00000160 0b 00 00 00 00 00 00 00 0b 00 00 00 00 00 00 00 |................| 00000170 00 74 6f 74 6f 2e 63 70 70 00 64 6f 5f 66 6f 6f |.test.cpp.do_foo| 00000180 00 00 00 00 00 00 00 00 10 00 00 00 00 00 00 00 |................| 00000190 04 00 00 00 03 00 00 00 fc ff ff ff ff ff ff ff |................| ``` As we can see, it only contains the ``do_foo`` string. Since ``foo`` is **a suffix** of ``do_foo``, ``st_name`` can point to a different offset of the **same string**. In this layout ``Elf_Sym("foo").st_name`` points to the offset **0x17C** and ``Elf_Sym("foo").st_name`` points to **0x17A**. Consequently, instead of taking the space of ``len(call_foo) + 1 + len(foo) + 1``, it only takes ``len(call_foo) + 1`` The consequence of this optimization is that we can't naively push back the symbols names in the ``.dynstr`` section. Instead, we have to sort the symbols names such as this optimization can take place. In addition to this strings optimization, ELF object files (``.o``) generated by Clang share the same section for the names of the sections and for the symbols' names. ```bash $ readelf -hWS ./hello.o ELF Header: Magic: 7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00 [...] Number of section headers: 11 Section header string table index: 1 Section Headers: [Nr] Name Type Address Off Size ES Flg Lk Inf Al [ 0] NULL 0000000000000000 000000 000000 00 0 0 0 [ 1] .strtab STRTAB 0000000000000000 000199 000078 00 0 0 1 [..] [10] .symtab SYMTAB 0000000000000000 0000c0 000090 18 1 4 8 ``` As we can notice, the *Section header string table index* of the ELF header indexes the ``.strtab`` which is also the section associated with the symbols' names (cf. the *link* attribute of the ``.symtab``). It results that we have to consider this kind of ELF file differently from regular libraries or executables. There are other nasty tricks like the management of the ELF constructors between Linux and Android but this will be covered in another blog post. ## The New ELF Builder For the historical context, I created LIEF during my internship at [Quarkslab](https://quarkslab.com/) with the supervision of [Serge-Sans-Paille](http://serge.liyun.free.fr) and [Adrien Guinet](https://aguinet.github.io/) and the trust/boost from [Fred Raynal](https://quarkslab.com/about/). Even though I had the chance to get valuable feedback and review from them, I clearly made poor design decisions in LIEF and the implementation of the ELF builder is one of them. Basically, the implementation is **recursive** such as in the extreme cases the builder re-computes the same information several times. In the new implementation, we added a new stage in the build process that pre-computes the offsets of the new sections and the data that need to be relocated. This pre-computation enables to know exactly which parts of the ELF structures need to be relocated according to the user's changes. This computation is managed by the [Layout](https://github.com/lief-project/LIEF/blob/2ae5327e86f50fe87733d8641d4e7bc3774e3087/src/ELF/ExeLayout.hpp) class which has two implementations depending on whether it is an ELF object or a library/executable. Compared to the previous ELF builder, this new implementation produces smaller files (with fewer ELF segments) as exposed in the following figure. This figure compares the number of segments between the former and the new implementation: ![Comparison of the number of segments generated](https://lief.re/blog/2022-01-23-new-elf-builder/bench_nb_segments.png) In addition, it supports larger binaries faster as a consequence of the new linear implementation of the ELF builder :) ![Comparison of the number of segments generated](https://lief.re/blog/2022-01-23-new-elf-builder/bench_time.png) To perform these benchmarks, we generated ELF binaries with the modifications described in the following script: ```python import lief elf: lief.ELF.Binary = lief.parse(file_path.as_posix()) # Force relocating the .dynamic/.dynstr elf.add_library("a_very_long_name.so") # For relocating the interpreter elf.interpreter = "/a/very/longlonglong/interpreter-1.2.3.bin" # Force relocating .dynsym / .gnu.hash table for i in range(10): elf.add_exported_function(0xdeadc0de + i, f"new_export_{i}") # Add a segment segment = lief.ELF.Segment() segment.type = lief.ELF.SEGMENT_TYPES.LOAD segment.content = [0xcc] * 0x23 elf.add(segment) elf.write("/tmp/bench.bin") ``` The raw results of the benchmark are also available [here](https://lief.re/blog/2022-01-23-new-elf-builder/benchmark.txt) ## Final Words These new improvements introduce breaking changes in the ELF binaries generated by LIEF but: 1. The final binary size should be smaller 2. The building time should be much faster We tried to cover most of the cases in the tests suite but some corner cases with exotic compilers or linkers might break the final binaries. Since this improvement aims at being in the next release, feel free to drop an email or to open an issue if you find a bug with this new implementation. --- *Since September I continue maintaining LIEF exclusively in my spare so issues and new features are addressed with more delay.* --- # Profiling C++ code with Frida (2nd part) > A practical follow-up on profiling C++ with Frida, including static-library hooks and the limitations introduced by different C++ ABIs. - Canonical URL: https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/ - Markdown: https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/index.md - Authors: Romain Thomas - Published: 2021-04-08T00:00:00Z - Modified: 2021-04-08T00:00:00Z - Tags: Frida, C++, profiling, ABI, Windows **Tl;DR** This blog post is not, strictly speaking, related to LIEF but it aims at completing the previous blog about profiling code with Frida. In particular, it exposes the limits of our approach regarding the Microsoft/Itanium ABI. Long story short, the previous code **does not work** on Linux/OSX for virtual functions. The previous blog post tried to show a use case of Frida to profile C++ functions. In particular, it exposed what we called a *trick* to convert a C++ member function into a ``void*``: ```cpp template inline void* cast_func(Func f) { union { Func func; void* p; }; func = f; return p; } ``` First, and as noticed by Julien Jorge, writing a union's field and accessing another field of this union is [undefined behavior](https://en.cppreference.com/w/cpp/language/union#Explanation): > It's undefined behavior to read from the member of the union that wasn't most recently written. > Many compilers implement, as a non-standard language extension, the ability to read inactive members of a union. Thanks also to the feedback from Julien Jorge, there is another issue when converting a C++ member function into a raw pointer. Basically, a member function pointer is not the same kind of pointer as a regular C function. While the regular size of a C function pointer is the same as ``sizeof(void*)``, the size of a member function pointer is usually greater: ```cpp struct Foo { void bar() {} }; int main() { printf("sizeof(&Foo::bar): %d\n", sizeof(&Foo::bar)); return 0; } ``` ```bash $ clang++ sizeof_member.cpp -o sizof_member $ ./sizeof_member sizeof(&Foo::bar): 16 ``` The layout of a member function pointer is ABI specific but according to LLVM's source code we can distinguish two ABI that describe this layout: 1. [Itanium CXX ABI](https://itanium-cxx-abi.github.io/cxx-abi/) which is used on Linux, iOS, OSX, Android, ... 2. Microsoft ## Itanium ABI For the Itanium CXX ABI and according to the official documentation, **non-virtual** functions have the following structure: ```cpp struct { uintptr_t ptr; ptrdiff_t adj; }; ``` Where, ``ptr`` is the address of the function and ``adj`` is an offset applied on ``this`` in the case of multi-inheritance. So in our *bad-coded* casting function ``cast_func()``, it works as expected for non-virtual functions since we access the first field ``ptr`` which is the function pointer. We can observe these two fields with the following piece of code [^1]: ```cpp template void print(Func f) { union { Func fcn; struct { uintptr_t ptr; ptrdiff_t adj; }; }; fcn = f; printf("%016lx | %016lx\n", ptr, adj); } ``` This outputs values such as: ```cpp struct Foo { void bar() {} }; int main() { print(&Foo::bar); return 0; } ``` ``` $ ./show_fields 00005568e19021e0 | 0000000000000000 ``` If ``bar()`` were a **virtual function**, the meaning of the ``ptr`` field would be different. Still according to the Itanium CXX ABI, the value of ``ptr`` in the case of a virtual function is 1 plus the offset of the function within the v-table. In particular, we can't access the address of the function without ``this`` since the vtable is embedded in the layout of the object. [^2] ![Itanium CXX ABI](https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/cxxabi.png) ## Microsoft ABI Regarding the Microsoft ABI, there is not as much documentation compared to the "Linux/OSX" ABI. LLVM supports this ABI as described in [clang/lib/CodeGen/MicrosoftCXXABI.cpp](https://github.com/llvm/llvm-project/blob/a59665930b87d7510002dcf1f292b290673a47d3/clang/lib/CodeGen/MicrosoftCXXABI.cpp) but I was still curious to know how (without LLVM) the layout of a function member pointer looks like. One could look at ``c1xx.dll/c2.dll`` located in the Visual Studio directory but these libraries are not straightforward to reverse. Alternately, we can try to infer the layout from the assembly code output. First of all, the result of ``sizeof()`` applied to a function member pointer is 16. 16 being twice a pointer's size on an 64-bits architecture, we can start following the Itanium ABI and confirm or infirm our choices: ```cpp struct FuncMemPtr { uintptr_t unknown1; uintptr_t unknown2; }; ``` Then we can *unpack* the fields of the function member pointer with the union trick: ```cpp struct Base1 { virtual void f() { } }; struct Base2 { virtual void g() {} }; struct Derived2 : Base2, Base1 { virtual void f() {} virtual void g() {} virtual h() {} }; template void info(Func f) { union { Func fcn; struct { uintptr_t unknown1; uintptr_t unknown2; }; }; fcn = f; } int main() { info(&Derived2::h); info(&Derived2::f); return 0; } ``` The layout of the non-virtual function ``Derived2::h()`` seems to follow the same layout as the Itanium ABI where we find the function pointer in the first field. ![Non virtual function layout](https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/msvc_non_virtual.png) For the **virtual function** ``Derived2::f``, we can notice a first memory write that fills the first field with a pointer to a thunk [^3] function while the second field contains a constant which matches the value of *this adjustor*. For the second field (*this adjustor*), we can switch from ``&Derived2::f`` to ``&Derived2::g`` to confirm that it changes accordingly to the output of ``/d1reportAllClassLayout`` ![Virtual function layout](https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/msvc_virtual.png) This leads to the following guessing: ```cpp struct MsvcCXXFuncMember { uintptr_t fnc_ptr; // That can be a thunk for virtual function int adjustor; // int because of mov DWORD and not mov QWORD in this assembly output }; ``` These two fields follow the [LLVM implementation](https://github.com/llvm/llvm-project/blob/4708a05da03038271a1a2c1cbdfe78aebfaa7afc/clang/lib/AST/MicrosoftCXXABI.cpp#L223-L234): ```cpp struct { // A pointer to the member function to call. If the member function is // virtual, this will be a thunk that forwards to the appropriate vftable // slot. void *FunctionPointerOrVirtualThunk; // An offset to add to the address of the vbtable pointer after // (possibly) selecting the virtual base but before resolving and calling // the function. // Only needed if the class has any virtual bases or bases at a non-zero // offset. int NonVirtualBaseAdjustment; // The offset of the vb-table pointer within the object. Only needed for // incomplete types. int VBPtrOffset; // An offset within the vb-table that selects the virtual base containing // the member. Loading from this offset produces a new offset that is // added to the address of the vb-table pointer to produce the base. int VirtualBaseAdjustmentOffset; }; ``` From LLVM, we also learn that the full layout can contain up to *four fields*. We can trigger the third field with the following change: ```diff {style=pastie} @@ -11,3 +11,3 @@ -struct Derived2 : Base2, Base1 { +struct Derived2 : Base2, virtual Base1 { virtual void f() {} @@ -32 +32,2 @@ } ``` ![VBPtrOffset](https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/msvc_third_field.png) The fourth field is a bit more tricky to trigger and the following code comes from the [LLVM test suite](https://github.com/llvm/llvm-project/blob/4fffbc150cca1638051b8ad2a20f4b8240df0869/clang/test/CodeGenCXX/microsoft-abi-member-pointers.cpp) [^4] ```cpp struct B1 { void foo(); int b; }; struct B2 { int b2; int v; void foo(); }; struct UnspecWithVBPtr; int UnspecWithVBPtr::*forceUnspecWithVBPtr; struct UnspecWithVBPtr : B1, virtual B2 { int u; void foo(); }; ``` ![VBPtrOffset](https://lief.re/blog/2021-04-08-profiling-cpp-code-with-frida-part2/msvc_fourth_field_e.png) We can notice that the result of ``sizeof()`` applied to ``UnspecWithVBPtr::foo`` is **24**: ``sizeof(uintptr_t) + 3 * sizeof(int) + padding`` ## Conclusion The profiler described in the first blog post works as expected for **non-virtual** but does not work with virtual functions that follow the Itanium ABI. To work with virtual functions, we would need to pass an extra parameter to the object that implements the virtual functions. By assuming that the vtable is placed at the beginning of the object's layout, we can support such functions with the following modifications: ```diff {style=pastie} diff --git a/main.cpp b/main.cpp index d30a0c1..65d18eb 100644 --- a/main.cpp +++ b/main.cpp +struct Foo { + virtual void bar() { + std::cout << "In bar" << std::endl; + } + uint8_t x = 1; +}; + @@ -88,9 +95,10 @@ struct Profiler { - void setup() { - PROFILE(LIEF::ELF::Parser::init); - PROFILE(LIEF::ELF::Parser::parse_segments); + template + void setup(const T& obj) { + const uintptr_t vtable = *reinterpret_cast(&obj); + profile_func(&Foo::bar, "Foo:bar", vtable); } @@ -98,8 +106,13 @@ struct Profiler { template - void profile_func(Func func, std::string name) { + void profile_func(Func func, std::string name, uintptr_t vtable = 0) { void* addr = cast_func(func); + + if (vtable > 0) { + const uintptr_t voff = reinterpret_cast(addr) - 1; + addr = *reinterpret_cast(vtable + voff); + } funcs[reinterpret_cast(addr)] = std::move(name); gum_interceptor_begin_transaction (ctx_->interceptor); gum_interceptor_attach (ctx_->interceptor, @@ -130,8 +143,9 @@ int main(int argc, const char** argv) { return 1; } + Foo f; Profiler& prof = Profiler::get(); - prof.setup(); - LIEF::ELF::Parser::parse(argv[1]); + prof.setup(f); + f.bar(); return 0; } ``` The Microsoft C++ ABI is poorly documented but the LLVM project is a good reference for that. One might also be interested in this presentation ([Bringing Clang and LLVM to Visual C++ users](https://llvm.org/devmtg/2013-11/slides/Kleckner-ClangVisualC++.pdf)) that outlines the challenges for LLVM developers to support this ABI. ## Acknowledgment Thanks to Julien Jorge for proofreading this post and his valuable feedback. [^1]: Which is still UB [^2]: ARM is an exception. In the 32-bit ARM representation, the `this` adjustment stored in `adj` is left-shifted by one. > The low bit of `adj` indicates whether `ptr` is a function pointer (including null) or the offset of a v-table entry. > A virtual member function pointer sets `ptr` to the v-table entry offset as if by ``reinterpret_cast(uintfnptr_t(offset))``. > A null member function pointer sets `ptr` to a null function pointer and must ensure that the low bit of `adj` is clear; > the upper bits of `adj` remain unspecified. [^3]: A thunk function is generated by the compiler as a *trampoline* to the right virtual function. This trampoline can also be used to fix ``this`` pointer with the given adjustor. [^4]: The layout of this code goes beyond my understanding --- # Profiling C++ code with Frida > Profile C/C++ code with Frida: hook functions through the frida-gum C API and use LIEF to enumerate the functions to instrument, without recompiling the target. - Canonical URL: https://lief.re/blog/2021-03-10-profiling-cpp-code-with-frida/ - Markdown: https://lief.re/blog/2021-03-10-profiling-cpp-code-with-frida/index.md - Authors: Romain Thomas - Published: 2021-03-10T00:00:00Z - Modified: 2021-03-10T00:00:00Z - Tags: Frida, C++, profiling Frida is a well-known reverse engineering framework that enables (along with other functionalities) to hook functions on closed-source binaries. While hooking is generally used to get dynamic information about functions for which we don't have the source code, this blog post introduces another use case to profile C/C++ code. ## Code Profiling LIEF starts to be quite mature but there are still some concerns regarding: 1. The speed (especially when rebuilding large ELF binaries) 2. The memory consumption 3. Compilation time These limitations are "quite" acceptable on modern computers but when we target embedded systems like iPhone or Android devices, it starts to reach the limits. Since (spoiler) I started to implement a parser for the Dyld shared cache and for parsing in-memory Mach-O files, I faced some of these issues. To address these problems, we must identify where are the bottleneck and ideally, without modifying too much the source code. To profile memory consumption, ``valgrind --tool=massif`` does the job pretty well out of the box: we don't need to pass extra compilation flags nor modifying the source code. Regarding the code execution, we can profile it with: 1. Valgrind (or QBDI?) 2. Inserting log functions in the source code 3. Using compiler instrumentation: ``-finstrument-functions`` In the context of profiling LIEF, I'm mostly interested in profiling the code at the functions level: "How long does a function take to be executed?" Valgrind provides CPU cycles that are somehow correlated to the execution time but it requires an extra processing step to identify the function's overhead. Moreover, since Valgrind instruments the code, it can take time to profile a large codebase. On the other hand, inserting log messages in the code is the easiest way to get the execution time of functions. I was not completely convinced with this solution since it adds log messages that are not always needed. Finally, Clang and GCC enable to instrument the source code through the ``-finstrument-functions`` compilation flag. This flag basically inserts the ``__cyg_profile_func_enter`` and ``__cyg_profile_func_exit`` functions at the beginning and at the end of the original functions. Frida works on compiled code and provides a mechanism (hook) to insert a callback before a given function and after the execution of the function. It is very similar to the ``-finstrument-functions``, except that it is done post-compilation. To set up a hook, we only have to provide a pointer to the function that aims at being hooked. In the context of profiling execution time, the callback at the beginning of the function can initialize a ``std::chrono`` object and the callback at the end of the function can print the time spent since the initialization of the ``std::chrono``. Let's take a simple example to explain what Frida does. If we have the following function: ```cpp void heavy_function() { for (size_t i = 0; i < 1000000; ++i) { // Code that takes time ... } } ``` Frida enables (from a logical point of view) to have: ```cpp void heavy_function() { frida_on_enter(); for (size_t i = 0; i < 1000000; ++i) { // Code that takes time ... } frida_on_leave(); } ``` ... without tweaking the compilation flags :) ## Frida Bootstrap Most documentation and blog posts about Frida use the JavaScript API. Frida also provides the frida-gum SDK [^1], which exposes a C API over the hook engine. This SDK comes with the ``frida-gum-example.c`` file, which shows how to set up the hook engine. Regarding the API of our profiler, we would like to have : ```cpp #include // Functions to profile profile(&LIEF::ELF::Parser::parse_symbol_version); profile(&LIEF::ELF::Parser::parse_segments); LIEF::ELF::Parser::parse("./sample.bin"); ``` And an output like: ```bash $ ./run LIEF::ELF::Parser::parse_symbol_version() took 39ms LIEF::ELF::Parser::parse_segments() took 109ms ``` I won't go through all the details of the implementation of the profiler since the source code is on [GitHub](https://github.com/lief-project/frida-profiler) but the next section covers some tricky parts. Firstly, and as mentioned previous section, Frida takes a **void\* pointer** on the function to hook. Therefore, we have to cast ``&LIEF::ELF::Parser::parse_symbol_version`` into a ``void*``. One might want to do `reinterpret_cast()` on the function pointer but it does not work. The trick here is to use a union to get the ``void*``: ```cpp template inline void* cast_func(Func f) { union { Func func; void* p; }; func = f; return p; } ``` Secondly, the example ``frida-gum-example.c`` uses an enum to identify the function being hooked: ```cpp typedef enum _ExampleHookId ExampleHookId; enum _ExampleHookId { EXAMPLE_HOOK_OPEN, EXAMPLE_HOOK_CLOSE }; ... gum_interceptor_attach (interceptor, GSIZE_TO_POINTER (gum_module_find_export_by_name (NULL, "open")), listener, GSIZE_TO_POINTER (EXAMPLE_HOOK_OPEN)); ``` In our case, we don't know beforehand which functions will be hooked or profiled by the user. Consequently, instead of using an enum we use the function's absolute address and we register its name in a map: ```cpp template void profile_func(Func func, std::string name) { void* addr = cast_func(func); funcs[reinterpret_cast(addr)] = std::move(name); gum_interceptor_begin_transaction(ctx_->interceptor); gum_interceptor_attach(ctx_->interceptor, /* Target */ reinterpret_cast(addr), /* Param */ reinterpret_cast(ctx_), /* id */ reinterpret_cast(addr)); gum_interceptor_end_transaction(ctx_->interceptor); } ``` Last but not least, we might want to profile private or protected functions. To enable the access to the Profiler to protected/private members we can *friend* an opaque Profile structure: ```cpp struct Profiler; namespace LIEF { class LIEF_API Parser : public LIEF::Parser { public: friend struct ::Profiler; ... }; } ``` ## Conclusion Through this blog post, we have shown that Frida also has some applications in the field of software engineering not only for reverse-engineering :) This approach can be quite convenient to isolate the profiling process from the compilation process. It also enables to quickly switch from a given SDK version to another as long as the profiled functions still exist. ```bash $ clang++ [-other-flags] LIEF-0.9.0/lib/libLIEF.a profile.cpp $ clang++ [-other-flags] LIEF-0.12.0/lib/libLIEF.a profile.cpp ``` The source code used in this blog post is available on GitHub: [lief-project/frida-profiler](https://github.com/lief-project/frida-profiler) [^1]: See ``frida-gum-devkit-14.2.13-linux-x86_64.tar.xz`` on https://github.com/frida/frida/releases --- # LIEF - Release 0.11.1 > LIEF 0.11.1 fixes PE Authentihash computation: section name handling, data directory coverage, and the return value of verify_signature(). - Canonical URL: https://lief.re/blog/2021-02-22-lief-0-11-1/ - Markdown: https://lief.re/blog/2021-02-22-lief-0-11-1/index.md - Authors: Romain Thomas - Published: 2021-02-22T00:00:00Z - Modified: 2021-02-22T00:00:00Z - Tags: release, PE, Authenticode **Tl;DR** LIEF v0.11.1 fixes some issues related to PE Authentihash computation. The new packages are available on PyPI and the SDKs can be downloaded on the official [website](https://lief.quarkslab.com/download/). Enjoy! LIEF 0.11.0 missed handling some cases in the processing of the PE Authentihash. This new release addresses these issues and the following blog post explains the cases we did not handle. ## Section name PE section's names are stored in a **fixed** char array (8 bytes) which means that a section's name can contain trailing bytes after the null char: ```cpp struct pe_section { char name[8]; uint32_t RVA; // ... }; ``` Before v0.11.1, LIEF didn't take into account the trailing bytes and stopped to read the section's name on the first null char: ```cpp this->name_ = std::string(header->name, sizeof(header->name)).c_str(); ``` This implementation has two drawbacks. First, we lose information since we don't store the extra trailing bytes. Regular binaries have zero trailing bytes after the first null char but some of them might use this spot to hide data. ![Section name with trailing bytes](https://lief.re/blog/2021-02-22-lief-0-11-1/section_table_e.png) Secondly, the **full** section name (i.e the whole 8 bytes) is used to compute the Authentihash. Therefore, if the first null char is followed by trailing bytes different from zero, the computed hash is inconsistent. ## Data directory According to the PE specifications [^1] the last entry of the data directory table must contain a null entry (i.e. an entry with an RVA and size set to 0). ![PE specifications require a last zero entry](https://lief.re/blog/2021-02-22-lief-0-11-1/pe_doc_e.png) It turns out that this requirement is not enforced by the loader. In the case of the binary ([bc203f2b6a...](https://www.virustotal.com/gui/file/bc203f2b6a928f1457e9ca99456747bcb7adbbfff789d1c47e9479aac11598af/detection)) the last entry is set to ``0x02b7bc68/0x01a7a0`` (used for watermarking?). ![Last data directory with non-zero entry](https://lief.re/blog/2021-02-22-lief-0-11-1/data_directory_e.png) In the previous versions of LIEF we assumed that the last entry of the data directory table was always zero. Since the last entry is used to compute the Authentihash value, it led to a bad signature while it was effectively correct. This issue has been addressed in the commit [3c65ffe](https://github.com/lief-project/LIEF/commit/3c65ffe2d65f0c6fe63e683e6deef41de2f395b1) ## Return value of ``verify_signature()`` As noticed by [Cedric Halbronn](https://twitter.com/saidelike) in the issue [issues/532](https://github.com/lief-project/LIEF/issues/532), the return value of [LIEF::PE::Binary::verify_signature](https://github.com/lief-project/LIEF/blob/f58605f94c365b5aedf75081ae9b0aebafd5cece/include/LIEF/PE/Binary.hpp#L147-L167) lacks information when the verification failed. The return value was either ``VERIFICATION_FLAGS.OK`` or ``VERIFICATION_FLAGS.BAD_SIGNATURE`` because of a *fail-fast* implementation of the verification flag. The function now returns flags as follows: ``` VERIFICATION_FLAGS.BAD_DIGEST | VERIFICATION_FLAGS.BAD_SIGNATURE | VERIFICATION_FLAGS.CERT_EXPIRED ``` ## Other issues One of the critical issues raised by [imidoriya](https://github.com/imidoriya) and fixed in the new version is the processing of the overlay data when the "data directory signature" is located in this area (c.f. [463bb0ec3...](https://www.virustotal.com/gui/file/463bb0ec399af716b9ec984dbc96590e180921af073df414b85a2a3b7c27516a/detection)). This kind of layout triggers a memory error on this part of the processing [Binary.cpp#L1174-L1187](https://github.com/lief-project/LIEF/blob/f58605f94c365b5aedf75081ae9b0aebafd5cece/src/PE/Binary.cpp#L1174-L1187). It has been addressed in the commit [05103f5](https://github.com/lief-project/LIEF/commit/05103f55a6cb993cb20735da3c7a6333e4f600e3) ## Acknowledgment Thanks to [Andrew Williams](https://twitter.com/SmugYeti) for providing the different samples that raised some of these errors! Thank you also to [Cedric Halbronn]() and the [CERT Gouvernemental of Luxembourg](https://www.govcert.lu/en/) for their feedback about the API. [^1]: https://docs.microsoft.com/en-us/windows/win32/debug/pe-format#optional-header-data-directories-image-only --- # LIEF - Release 0.11.0 > LIEF 0.11.0: refactored PE Authenticode parsing with signature verification, pefile-compatible imphash, a faster ELF builder, and Ninja-based Windows CI. - Canonical URL: https://lief.re/blog/2021-01-19-lief-0-11-0/ - Markdown: https://lief.re/blog/2021-01-19-lief-0-11-0/index.md - Authors: Romain Thomas - Published: 2021-01-19T00:00:00Z - Modified: 2021-01-19T00:00:00Z - Tags: release, PE, Authenticode, ELF, Windows **Tl;DR** LIEF v0.11.0 is out. The main changelog is available [here](https://lief.quarkslab.com/doc/stable/changelog.html#v0.11.0) and packages can be downloaded on the [official website](https://lief.quarkslab.com/download). ## Installation As for the previous versions, release packages are available on the [GitHub release page](https://github.com/lief-project/LIEF/releases/tag/0.11.0) and Python packages can be installed from PyPI: ```bash $ pip install [--user] lief==0.11.0 ``` ## Release Highlight It has spent more than one year since the release of the version [0.10.1](https://lief.quarkslab.com/doc/latest/changelog.html#november-29-2019) but we are glad to announce that **LIEF v0.11.0** is finally out! This new version does not introduce a lot of new features but rather small improvements in the different formats. One of the main changes in terms of new functionalities is the refactoring of the PE Authenticode. We fixed parsing issues and we implemented verification functions so that we can now verify a PE signed binary through: ```python import lief pe = lief.parse("signed.exe") assert pe.verify_signature() == lief.PE.Signature.VERIFICATION_FLAGS.OK ``` We also improved the computation of *imphash* so that it can generate the same value as [pefile](https://github.com/erocarrera/pefile) (and therefore, Virus Total) ```python pe = lief.parse("example.exe") vt_imphash = lief.PE.get_imphash(pe, lief.PE.IMPHASH_MODE.PEFILE) lief_imphash = lief.PE.get_imphash(pe, lief.PE.IMPHASH_MODE.DEFAULT) ``` Regarding the contributions, [Janusz Lisiecki](https://github.com/JanuszL) fixed a performance issue in the **ELF builder** that moved from `N²` computations to `N log(N)`. His contribution raised a major weakness in LIEF: performances issue when re-building objects. We started to refactor the whole ELF builder to avoid recursive calls. [Adrien Guinet](https://github.com/aguinet) updated the [bin2lib tutorial](https://lief.quarkslab.com/doc/latest/tutorials/08_elf_bin2lib.html#warning-for-glibc-2-29-users) to support recent versions of glibc, which introduced the [`DF_1_PIE`](https://lief.quarkslab.com/doc/latest/api/python/elf.html#lief.ELF.DYNAMIC_FLAGS_1.PIE) flag. [kohnakagawa](https://github.com/kohnakagawa) and [Clcanny](https://github.com/Clcanny) also fixed various issues related to the ELF & PE formats. ## Ninja on Windows & CI We improved AppVeyor Windows CI to be more efficient on the compiler cache. It results in a decrease of 1-hour compilation time to ~20 minutes thanks to [sccache](https://github.com/mozilla/sccache) and Ninja. If Ninja is installed on Windows, one can now use the ``--ninja`` flag when calling ``setup.py``: ```text $ python.exe .\setup.py --ninja build install [--user] ``` Using Ninja on Windows requires to invoke the ``vcvarsall.bat`` script beforehand. This script can be tricky to locate depending on the MSVC versions. Thankfully, setuptools provides the [msvc.msvc14_get_vc_env()](https://github.com/pypa/setuptools/blob/6ad2fb0b78d11e22672f56ef9d65d13ebd3475a9/setuptools/msvc.py#L293) helper to get the environment variables that need to populate the calling script. We use it in LIEF's ``setup.py`` as follows: ```python ... env = os.environ if platform.system() == "Windows": from setuptools import msvc if build_with_ninja: arch = 'x64' if is64 else 'x86' ninja_env = msvc.msvc14_get_vc_env(arch) env.update(ninja_env) else: ... ... ``` Regarding the CI, we added Android and iOS SDK packages as well as Python wheels for Linux AArch64 (``manylinux2014`` compliant). The nightly builds are available on the [gh-pages](https://github.com/lief-project/packages/tree/gh-pages) branch of the repository [lief-project/packages](https://github.com/lief-project/packages): - The [**sdk**](https://github.com/lief-project/packages/tree/gh-pages/sdk) directory contains a shared and a static version of LIEF library for iOS, macOS, Android, Windows, Linux, ... - The [**lief**](https://github.com/lief-project/packages/tree/gh-pages/lief) directory contains the Python wheels for the supported platforms ## What's next We have a few ideas of what would like to improve and introduce in the next releases of LIEF which includes: - Refactoring the ELF builder to address performances issues (see also [#482](https://github.com/lief-project/LIEF/issues/482)) - Supporting OAT/VDEX/CDEX for Android 9, 10 and 11 - Supporting Mach-O signature (as for PE Authenticode) - Supporting Android packed relocations (in the parser and in the builder) - Improving the C API to ease Rust bindings - Supporting DART snapshot formats to ease reverse-engineering of Flutter applications. *Spoiler: we can process all the clusters of a snapshot for a fixed version of the DART runtime.* - `+=` Fixing issues Although LIEF's plans mostly follow Quarkslab's needs, the R&D time we have, and the topics we enjoy working on, we are open to the development of private or public features as it has been done for improving PE Authenticode. ## Acknowledgment Thank you to [CERT Gouvernemental of Luxembourg](https://www.govcert.lu/en/) that sponsored new functionalities in this release. Thanks also to [Quarkslab](https://www.quarkslab.com) for the time allocated to make this release. [ ![Logo Quarkslab](https://lief.re/blog/2021-01-19-lief-0-11-0/logo-quarkslab.png) ](https://www.quarkslab.com) [ ![Logo CERT Gouvernemental Luxembourg](https://lief.re/blog/2021-01-19-lief-0-11-0/logo-govcert-lu.png) ](https://www.govcert.lu/en/) --- # LIEF - Release 0.9.0 > Major changes in LIEF 0.9, plus work-in-progress features planned for future releases. - Canonical URL: https://lief.re/blog/2018-06-11-lief-0-9-0/ - Markdown: https://lief.re/blog/2018-06-11-lief-0-9-0/index.md - Authors: Romain Thomas - Published: 2018-06-11T00:00:00Z - Modified: 2018-06-11T00:00:00Z - Tags: release, Android, JSON ## Installation Release packages are available on the [GitHub page](https://github.com/lief-project/LIEF/releases/tag/0.9.0) and Python package can be installed with: ```bash $ pip install [--user] lief==0.9.0 ``` ## Release highlight ### Android Formats This new version of LIEF comes with support for Android formats related to the ART runtime: OAT, VDEX, DEX and ART. As the OAT format is a derivation of ELF, it made sense to add it in LIEF. Basically, this format is used by Android to wrap native code being the result of Dalvik bytecode optimization. Regarding VDEX, DEX, and ART, these formats have somehow a relation with OAT and therefore we also choose to add them. For more information about these Android formats and how to use them, a tutorial is available in the LIEF documentation: [Android Formats](https://lief.quarkslab.com/doc/stable/tutorials/10_android_formats.html). We can currently only parse these formats, but support for modification will be added incrementally. Some attacks rely on modifying the OAT format, as Collin Mulliner explains in "*Inside Android’s SafetyNetAttestation: Attack and Defense*" [^1]. Tencent’s Xuanwu Lab also discusses them in "*How Samsung Secures Your Wallet & How To Break It*" [^2]. In a future version, we plan to provide an API for adding native code to OAT. ### JSON serialization As one purpose of this project is to provide an API that can be easily integrated in other projects, we are glad to announce that JSON serialization is now available for **all** LIEF objects. It means that one can now access to format information through a JSON interface. Previous versions had a JSON support for ELF and PE formats, the v0.9 now supports all formats and all objects. Objects can be serialized with the ``lief.to_json`` function: ```python import lief gcc = lief.parse("/usr/bin/gcc") lief.to_json(gcc.header) { 'entrypoint': 4209824, 'file_type': 'EXECUTABLE', 'header_size': 64, 'identity_class': 'CLASS64', 'identity_data': 'LSB' } libSystem = lief.parse("/usr/lib/libSystem.dylib") lief.to_json(libSystem.commands[1]) { 'command': 'SEGMENT', 'command_offset': 492, 'command_size': 464, 'content_hash': 18446744072658165641, 'data_hash': 1841536728, 'file_offset': 8192, 'file_size': 4096, 'flags': 0, 'init_protection': 3, 'max_protection': 7, 'name': '__DATA', 'numberof_sections': 6, 'sections': ['__nl_symbol_ptr', '__la_symbol_ptr', '__mod_init_func', '__const', '__data', '__common'], 'virtual_address': 8192, 'virtual_size': 4096 } ``` One can also disable the JSON module using a CMake configuration flag: ```bash $ cmake -DLIEF_ENABLE_JSON=off ... ``` ## What's next LIEF v0.9 still has a poor support for Mach-O modification and only supports modifications on header and some Load commands. One of the primitives to do more general modification on Mach-O format is the ability to add arbitrary Load commands. Some tools [^3] [^4] exist to add commands, but they usually use padding between the load command table and the raw content or they remove / replace existing one. The main limitation with this technique is that the number of load command which can be added depends on the size of the padding. In LIEF, we took advantage of the fact that Mach-O are PIE to *shift* the content that follow the load command table. This enable us to inject more than one or two commands. To keep a consistent state of format (relocations, segment's virtual address, ...), the Mach-O builder of LIEF rebuilds the export-trie, regenerates binding opcode, rebase opcodes, ... In our tests, we succeeded in adding arbitrary number of ``LC_DYLIB`` command in clang as well as adding 10 new sections in the ``__TEXT`` segment. We are currently working on stabilization of the instrumentation process, but it should be merged soon in then master branch. Stay tuned! We will be also be presenting about file formats instrumentation at [Recon Montréal](https://recon.cx/2018/montreal/) and [Pass The Salt](https://2018.pass-the-salt.org/programme/#instrumentation) for a talk about file formats instrumentation. In this talk we will present techniques to perform code injection, hooking by using formats. [^1]: Slide 58 of [Inside SafetyNet Attestation Attacks and Defense](https://www.mulliner.org/collin/publications/inside_safetynet_attestation_attacks_and_defense_mulliner2017_ekoparty.pdf). [^2]: Slide 89 of [How Samsung Secures Your Wallet And How To Break It](https://www.blackhat.com/docs/eu-17/materials/eu-17-Ma-How-Samsung-Secures-Your-Wallet-And-How-To-Break-It.pdf). [^3]: [insert_dylib](https://github.com/Tyilo/insert_dylib). [^4]: [optool](https://github.com/alexzielenski/optool) --- # Have fun with LIEF and Executable Formats! > LIEF 0.8.3: libFuzzer integration, Dockerlief, ELF symbol renaming and hiding, DT_RUNPATH injection, PE load configuration, and Mach-O dyld info parsing. - Canonical URL: https://lief.re/blog/2017-10-30-lief-0-8-3/ - Markdown: https://lief.re/blog/2017-10-30-lief-0-8-3/index.md - Authors: Romain Thomas - Published: 2017-10-30T00:00:00Z - Modified: 2017-10-30T00:00:00Z - Tags: release, ELF, PE, Mach-O, fuzzing, malware-analysis **Tl;DR** LIEF v0.8.3 is out. The main changelog is available [here](https://lief.quarkslab.com/doc/stable/changelog.html#october-16-2017) and packages can be downloaded on the [official website](https://lief.quarkslab.com/download). ## Development process We attach a great importance to the automation of some development tasks like testing, distributing, packaging, etc. Here is a summary of these processes: Each commits is tested on * Linux - x86-64 - Python{2.7, 3.5, 3.6} * Windows - x86 / x86-64 - Python{2.7, 3.5, 3.6} * OSX - x86-64 - Python{2.7, 3.5, 3.6} The test suite includes: * Tests on the Python API * Tests on the C API * Tests on the parsers * Tests on the builders If tests succeeds packages are automatically uploaded on the https://github.com/lief-project/packages repository. For tagged version, packages are uploaded on the GitHub release page: https://github.com/lief-project/LIEF/releases. ## Dockerlief To facilitate the compilation and the use of LIEF, we created the [Dockerlief](https://github.com/lief-project/Dockerlief) repo which includes various [Dockerfiles](https://github.com/lief-project/Dockerlief/tree/v0.1.0/dockerlief/dockerfiles) as well as the ``dockerlief`` utility. ``dockerlief`` is basically a wrapper on *docker build* . Among Dockerfiles, we provide a [Dockerfile](https://github.com/lief-project/Dockerlief/blob/v0.1.0/dockerlief/dockerfiles/android.docker) to cross compile LIEF for Android (``ARM``, ``AARCH64``, ``x86``, ``x86-64``) To cross compile LIEF for Android ARM, one can run: ```bash $ dockerlief build --api-level 21 --arm lief-android [INFO] - Location of the Dockerfiles: ~/dockerfiles [INFO] - Building Dockerfile: 'lief-android' [INFO] - Target architecture: armeabi-v7a [INFO] - Target API Level: 21 ``` The SDK package ``LIEF-0.8.3-Android_API21_armeabi-v7a.tar.gz`` is automatically pulled from the Docker to the current directory. ## Integration of libFuzzer Fuzzing our own library is a good way to detect bugs, memory leak, unsanitized inputs ... Thus, we integrated [libFuzzer](https://llvm.org/docs/LibFuzzer.html) in the project. Fuzzing the LIEF ELF, PE, Mach-O parser is as simple as: ```cpp #include #include #include extern "C" int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) { std::vector raw = {data, data + size}; try { std::unique_ptr b{LIEF::Parser::parse(raw)}; } catch (const LIEF::exception& e) { std::cout << e.what() << std::endl; } return 0; } ``` To launch the fuzzer, one can run the following commands: ```bash $ make fuzz-elf # Launch ELF Fuzzer $ make fuzz-pe # Launch PE Fuzzer $ make fuzz-macho # Launch MachO Fuzzer $ make fuzz # Launch ELF, PE and MachO Fuzzer ``` ## ELF ### Play with ELF symbols - Part 2 In the [tutorial #03](https://lief.quarkslab.com/doc/stable/tutorials/03_elf_change_symbols.html) we demonstrated how to swap dynamic symbols between a binary and a library. In this part, we will see how we can rename these symbols. Changing symbol names is not a trivial modification, since modifying the string table of the ``PT_DYNAMIC`` segment has side effects: * It requires to update the hash table (GNU Hash / SYSV). * It usually requires to extend the ``DYNAMIC`` part of the ELF format. The previous version of LIEF already implements the rebuilding of the hash table but not the extending of the ``DYNAMIC`` part. With the ``v0.8.3`` we can extend the ``DYNAMIC`` part. Therefore: * We can add new entries in the ``.dynamic`` section * We can change dynamic symbols names * We can change ``DT_RUNPATH`` and ``DT_RPATH`` without length restriction We will rename all **imported** functions of ``gpg`` that are imported from ``libgcrypt.so.20`` into ``a_very_long_name_of_function_XX`` and all exported functions of ``libgcrypt.so.20`` into the same name (*XX* is the symbol index). [^1] ```python import lief # Load targets gpg = lief.parse("/usr/bin/gpg") libgcrypt = lief.parse("/usr/lib/libgcrypt.so.20") # Change names for idx, lsym in enumerate(filter(lambda e : e.exported, libgcrypt.dynamic_symbols)): new_name = 'a_very_long_name_of_function_{:d}'.format(idx) print("New name for '{}': {}".format(lsym.name, new_name)) for bsym in filter(lambda e : e.name == lsym.name, gpg.dynamic_symbols): bsym.name = new_name lsym.name = new_name # Write back binary.write(gpg.name) libgcrypt.write(libgcrypt.name) ``` By using ``readelf`` we can check that function names have been modified: ```bash $ readelf -s ./gpg|grep "a_very_long_name" 2: 0000000000000000 0 FUNC GLOBAL DEFAULT UND a_very_long_name_of_funct@GCRYPT_1.6 (2) 3: 0000000000000000 0 FUNC GLOBAL DEFAULT UND a_very_long_name_of_funct@GCRYPT_1.6 (2) 11: 0000000000000000 0 FUNC GLOBAL DEFAULT UND a_very_long_name_of_funct@GCRYPT_1.6 (2) 13: 0000000000000000 0 FUNC GLOBAL DEFAULT UND a_very_long_name_of_funct@GCRYPT_1.6 (2) ... $ readelf -s ./libgcrypt.so.20|grep "a_very_long_name" 88: 000000000000d050 6 FUNC GLOBAL DEFAULT 10 a_very_long_name_of_funct@@GCRYPT_1.6 89: 000000000000dcd0 69 FUNC GLOBAL DEFAULT 10 a_very_long_name_of_funct@@GCRYPT_1.6 90: 000000000000d310 34 FUNC GLOBAL DEFAULT 10 a_very_long_name_of_funct@@GCRYPT_1.6 91: 000000000000de70 81 FUNC GLOBAL DEFAULT 10 a_very_long_name_of_funct@@GCRYPT_1.6 ... ``` ![GPG Binary in IDA](https://lief.re/blog/2017-10-30-lief-0-8-3/ida_gpg.png) Now if we run the new ``gpg`` binary, we get the following error: ```bash $ ./gpg --output bar.txt --symmetric ./foo.txt relocation error: ./gpg: symbol a_very_long_name_of_function_8, version GCRYPT_1.6 not defined in file libgcrypt.so.20 with link time reference ``` Because the Linux loader tries to resolve the function ``a_very_long_name_of_function_8`` against ``/usr/lib/libgcrypt.so.20`` and that library doesn't include the updated names we get the error. One way to fix this error is to set the environment variable ``LD_LIBRARY_PATH`` to the current directory: ```bash $ LD_LIBRARY_PATH=. ./gpg --output bar.txt --symmetric ./foo.txt $ xxd ./bar.txt|head -n1 00000000: 8c0d 0407 0302 c5af 9fba cab1 9545 ebd2 .............E.. $ LD_LIBRARY_PATH=. ./gpg --output foo_decrypted.txt --decrypt ./bar.txt $ xxd ./foo_decrypted.txt|head -n1 00000000: 4865 6c6c 6f20 576f 726c 640a Hello World. ``` Another way to fix it is to add a new entry in ``.dynamic`` section. As mentioned at the beginning, we can now add new entries in the ``.dynamic`` so let's add a ``DT_RUNPATH`` entry with the ``$ORIGIN`` value so that the Linux loader resolves the modified ``libgcrypt.so.20`` instead of the system one: ```python ... # Add a DT_RUNPATH entry gpg += lief.ELF.DynamicEntryRunPath("$ORIGIN") # Write back binary.write(gpg.name) libgcrypt.write(libgcrypt.name) ``` And we don't need the ``LD_LIBRARY_PATH`` anymore: ```bash $ readelf -d ./gpg|grep RUNPATH 0x000000000000001d (RUNPATH) Library runpath: [$ORIGIN] $ ./gpg --decrypt ./bar.txt gpg: AES encrypted data gpg: encrypted with 1 passphrase Hello World ``` [^1]: All Python examples are done with the 3.5 version ### Hiding its symbols While IDA v7.0 has been released recently, among the [changelog](https://www.hex-rays.com/products/ida/7.0/index.shtml) one can notice two changes: * *ELF: describe symbols using symtab from DYNAMIC section* * *ELF: IDA now uses the PHT by default instead of the SHT to load segments from ELF files* These changes are partially true. Let's see what go wrong in IDA with the following snippet: ```python id = lief.parse("/usr/bin/id") dynsym = id.get_section(".dynsym") dynsym.entry_size = dynsym.size // 2 id.write("id_test") ``` This snippet defines the size of **one** symbol as the entire size of ``.dynsym`` section divided by 2. The *normal* size of ELF symbols would be: ```python >>> print(int(lief.ELF.ELF32.SIZES.SYM)) # For 32-bits 16 >>> print(int(lief.ELF.ELF64.SIZES.SYM)) # For 64-bits 24 ``` In the case of the 64-bits ``id`` binary, we set this size to **924**. When opening ``id_test`` in IDA and forcing to use **Segment** for parsing and not **Sections** we get the following *imports* : ![IDA Loading Options](https://lief.re/blog/2017-10-30-lief-0-8-3/ida_loading.png) ![IDA Imports overview](https://lief.re/blog/2017-10-30-lief-0-8-3/ida_imports.png) Only one import is resolved and the others are **hidden**. Note that ``id_test`` is still executable: ```bash $ id_test uid=1000(romain) gid=1000(romain) ... ``` By using ``readelf`` we can still retrieve the symbols and we have an error indicating that symbol size is corrupted. ```bash $ readelf -s id_test readelf: Error: Section 5 has invalid sh_entsize of 000000000000039c readelf: Error: (Using the expected size of 24 for the rest of this dump) Symbol table '.dynsym' contains 77 entries: Num: Value Size Type Bind Vis Ndx Name 0: 0000000000000000 0 NOTYPE LOCAL DEFAULT UND 1: 0000000000000000 0 FUNC GLOBAL DEFAULT UND endgrent@GLIBC_2.2.5 (2) 2: 0000000000000000 0 FUNC GLOBAL DEFAULT UND __uflow@GLIBC_2.2.5 (2) 3: 0000000000000000 0 FUNC GLOBAL DEFAULT UND getenv@GLIBC_2.2.5 (2) 4: 0000000000000000 0 FUNC GLOBAL DEFAULT UND free@GLIBC_2.2.5 (2) 5: 0000000000000000 0 FUNC GLOBAL DEFAULT UND abort@GLIBC_2.2.5 (2) ... ``` In LIEF the (dynamic) symbol table address is computed through the ``DT_SYMTAB`` from the ``PT_DYNAMIC`` segment. To compute the number of dynamic symbols LIEF uses three heuristics: 1. Based on hash tables ([Gnu Hash](https://github.com/lief-project/LIEF/blob/0.8.3/src/ELF/Parser.tcc#L711-L796) / [SYSV Hash](https://github.com/lief-project/LIEF/blob/0.8.3/src/ELF/Parser.tcc#L690-L708)) 2. Based on [relocations](https://github.com/lief-project/LIEF/blob/0.8.3/src/ELF/Parser.tcc#L513-L648) 3. Based on [sections](https://github.com/lief-project/LIEF/blob/0.8.3/src/ELF/Parser.tcc#L653-L672) Malwares start to use this kind of corruption as we will see in the next part. ### Rootnik Malware Rootnik is a malware targeting Android devices. It has been analyzed by Fortinet security researcher. A full analysis of the malware is available on the [Fortinet blog](https://blog.fortinet.com/2017/07/09/unmasking-android-malware-a-deep-dive-into-a-new-rootnik-variant-part-i). This part is focused on the ELF format analysis of one component: ``libshell``. Actually there are two libraries ``libshella_2.10.3.1.so`` and ``libshellx_2.10.3.1.so``. As they have the same purpose, we will use the x86 version. First if we look at the ELF sections of ``libshellx_2.10.3.1.so`` we can notice that the **address**, **offset**, and **size** of some sections like ``.text``, ``.init_array``, ``.dynstr``, ``.dynsym`` are set to 0. This kind of modification is used to disturb tools that rely on **sections** to parse some ELF structures (like objdump, readelf, IDA ...) ```bash $ readelf -S ./libshellx-2.10.3.1.so There are 21 section headers, starting at offset 0x2431c: Section Headers: [Nr] Name Type Addr Off Size ES Flg Lk Inf Al [ 0] NULL 00000000 000000 000000 00 0 0 0 [ 1] .dynsym DYNSYM 00000114 000114 000300 10 A 2 1 4 [ 2] .dynstr STRTAB 00000414 000414 0001e2 00 A 0 0 1 [ 3] .hash HASH 00000000 000000 000000 04 A 1 0 4 [ 4] .rel.dyn REL 00000000 000000 000000 08 A 1 0 4 [ 5] .rel.plt REL 00000000 000000 000000 08 AI 1 6 4 [ 6] .plt PROGBITS 00000000 000000 000000 04 AX 0 0 16 [ 7] .text PROGBITS 00000000 000000 000000 00 AX 0 0 16 [ 8] .code PROGBITS 00000000 000000 000000 00 AX 0 0 16 [ 9] .eh_frame PROGBITS 00000000 000000 000000 00 A 0 0 4 [10] .eh_frame_hdr PROGBITS 00000000 000000 000000 00 A 0 0 4 [11] .fini_array FINI_ARRAY 00000000 000000 000000 00 WA 0 0 4 [12] .init_array INIT_ARRAY 00000000 000000 000000 00 WA 0 0 4 [13] .dynamic DYNAMIC 0000ce50 00be50 0000f8 08 WA 2 0 4 [14] .got PROGBITS 00000000 000000 000000 00 WA 0 0 4 [15] .got.plt PROGBITS 00000000 000000 000000 00 WA 0 0 4 [16] .data PROGBITS 00000000 000000 000000 00 WA 0 0 16 [17] .bss NOBITS 0000d398 00c395 000000 00 WA 0 0 4 [18] .comment PROGBITS 00000000 00c395 000045 01 MS 0 0 1 [19] .note.gnu.gold-ve NOTE 00000000 00c3dc 00001c 00 0 0 4 [20] .shstrtab STRTAB 00000000 024268 0000b1 00 0 0 1 Key to Flags: W (write), A (alloc), X (execute), M (merge), S (strings), I (info), L (link order), O (extra OS processing required), G (group), T (TLS), C (compressed), x (unknown), o (OS specific), E (exclude), p (processor specific) ``` If we open the given library in IDA we have no exports, no imports, and no sections: ![Libshell in IDA](https://lief.re/blog/2017-10-30-lief-0-8-3/libshell_ida.png) Based on the segments and dynamic entries we can recover most of these information: * ``.init_array`` address and size are available through the ``DT_INIT_ARRAY`` and ``DT_INIT_ARRAYSZ`` entries * ``.dynstr`` address and size are available through the ``DT_STRTAB`` and ``DT_STRSZ`` * ``.dynsym`` address is available through the ``DT_SYMTAB`` The script [recover_shellx.py](https://gist.github.com/romainthomas/c262be2c29ab374451663b9f6dcdfc0d) recovers the missing values, patch sections and rebuild a *fixed* library. ![Libshell fixed with LIEF](https://lief.re/blog/2017-10-30-lief-0-8-3/libshell_FIXED_ida.png) Now if we open the new ``libshellx-2.10.3.1_FIXED.so`` we have access to imports / exports and some sections. The ``.init_array`` section contains 2 functions: * ``tencent652524168491435794009`` * ``sub_60C0`` The ``tencent652524168491435794009`` function basically do a stack alignment and the ``sub_60C0`` is **one** of the decryption routines [^3]. This function is obfuscated with graph flattening and looks like to O-LLVM graph flattening passe [^2]: ![libshell obfuscated CFG with O-LLVM](https://lief.re/blog/2017-10-30-lief-0-8-3/shellx_cfg.png) Fortunately, only a few *"relevant blocks"* remain, and they are not obfuscated. The function ``sub_60C0`` basically iterates over the program headers to find the encrypted one and decrypt it using a custom algorithm (based on shift, xor, etc). ![libshell cfg part 1](https://lief.re/blog/2017-10-30-lief-0-8-3/shellx_cfg_1.2.png) ![libshell cfg part 2](https://lief.re/blog/2017-10-30-lief-0-8-3/shellx_cfg_2.2.png) --- [^2]: As mentioned in the Fortinet blog post, the library is packed. [^3]: See the blog post about O-LLVM analysis: https://blog.quarkslab.com/deobfuscation-recovering-an-ollvm-protected-program.html ## Triggering CVE-2017-1000249 The [CVE-2017-1000249](http://seclists.org/oss-sec/2017/q3/397) is a stack based buffer overflow in the ``file`` utility. It affects the versions ``5.29``, ``5.30`` and ``5.31``. Basically the overflow occurs in the size of the note description. Using LIEF we can trigger the overflow as follows: ```python target = lief.parse("/usr/bin/id") note_build_id = target[lief.ELF.NOTE_TYPES.BUILD_ID] note_build_id.description = [0x41] * 30 target.write("id_overflow") ``` ```bash $ file --version file-5.29 magic file from /usr/share/file/misc/magic $ id_overflow uid=1000(romain) gid=1000(romain) ... $ file id_overflow *** buffer overflow detected ***: file terminated ./id_overflow: [1] 3418 abort (core dumped) file ./id_overflow ``` Here is the commit that introduced the bug: [9611f3](https://github.com/file/file/commit/9611f31313a93aa036389c5f3b15eea53510d4d1#diff-bc5c24ef9f39a5f4963ca28ecbc645b3L512) ## PE The *Load Configuration* directory is now parsed into the [LoadConfiguration](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/PE/LoadConfigurations/LoadConfiguration.hpp) object. This structure evolves with the Windows versions and LIEF has been designed to support this evolution. You can take a look at [LoadConfigurationV0](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/PE/LoadConfigurations/LoadConfigurationV0.hpp#L47-L52), [LoadConfigurationV6](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/PE/LoadConfigurations/LoadConfigurationV6.hpp#L47-L54). One can find the different versions of this structure in the following directories: * ``include/LIEF/PE/LoadConfigurations`` * ``src/PE/LoadConfigurations`` The current version of LIEF is able to parse the structure up to Windows 10 build 15002 with the *hotpatch table offset*. Here are some examples of the ``LoadConfiguration`` API: ```python >>> target = lief.parse("PE64_x86-64_binary_WinApp.exe") >>> target.has_configuration True >>> config = target.load_configuration >>> config.version WIN_VERSION.WIN10_0_15002 >>> hex(config.guard_rf_failure_routine) '0x140001040' ``` LIEF also provides an API to serialize any ELF or PE objects into JSON [^4] For examples to transform ``LoadConfiguration`` object into Json: ```python >>> from lief import to_json >>> to_json(config) '{"characteristics":248,"code_integrity":{"catalog":0,"catalog_offset":0 ... }}' # Not fully printed ``` One can also serialize the whole Binary object: ```python >>> to_json(target) '{"data_directories":[{"RVA":0,"size":0,"type":"EXPORT_TABLE"},{"RVA":62584,"section" ...}}' # # Not fully printed ``` [^4]: This feature is not yet available for MachO objects ## Mach-O For Mach-O binary, dynamic executables embed the ``LC_DYLD_INFO`` command which is associated with the ``dyld_info_command`` structure. The structure is basically a *list* of offsets and sizes pointing to other data structures. From ``/usr/lib/mach-o/loader.h`` the structure looks like this: ```cpp struct dyld_info_command { uint32_t cmd; uint32_t cmdsize; uint32_t rebase_off; uint32_t rebase_size; uint32_t bind_off; uint32_t bind_size; uint32_t weak_bind_off; uint32_t weak_bind_size; uint32_t lazy_bind_off; uint32_t lazy_bind_size; uint32_t export_off; uint32_t export_size; }; ``` The ``dyld`` loader uses this structure to: * Rebase the executable * Bind symbols to addresses * Retrieve exported functions (or symbols) Whereas in the ELF and PE format relocations are basically a **table**, Mach-O format uses **byte streams** to rebase the image and to bind symbols with addresses. For exports it uses a **trie** as subjacent structure. In the new version of LIEF, the Mach-O parser is able to handle these underlying structures to provide a user-friendly API: The export trie is represented by the [ExportInfo](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/MachO/ExportInfo.hpp) object which is usually tied to a [Symbol](https://github.com/lief-project/LIEF/blob/master/include/LIEF/MachO/Symbol.hpp). The binding byte stream is represented trough the [BindingInfo](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/MachO/BindingInfo.hpp) object. For the rebase byte stream, the parser create *virtual relocations* to model the rebasing process. These *virtual relocations* are represented by the [RelocationDyld](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/MachO/RelocationDyld.hpp) object and among other attributes it contains ``address``, ``size`` and ``type`` [^5]. Here is an example using the Python API: ```python >>> id = lief.parse("/usr/bin/id") >>> print(id.relocations[0]) 100002000 POINTER 64 DYLDINFO __DATA.__eh_frame dyld_stub_binder >>> print(id.has_dyld_info) True >>> dyldinfo = id.dyld_info >>> print(dyldinfo.bindings[0]) Class: STANDARD Type: POINTER Address: 0x100002010 Symbol: ___stderrp Segment: __DATA Library: /usr/lib/libSystem.B.dylib >>> print(dyldinfo.exports[0]) Node Offset: 18 Flags: 0 Address: 0 Symbol: __mh_execute_header ``` [^5]: Due to the inheritance relationship and abstraction these attributes are located in the [MachO::Relocation](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/MachO/Relocation.hpp#L72-L83) and [LIEF::Relocation](https://github.com/lief-project/LIEF/blob/0.8.3/include/LIEF/Abstract/Relocation.hpp#L39-L43) objects. ## Conclusion In this release we did a large improvement of the ELF builder. Mach-O and PE parts gain new objects and new functions. LIEF is now available on [PyPI](https://pypi.python.org/pypi/lief) and can be added in the *requirements* of Python projects whatever the Python version and the target platform. Since the ``v0.7.0`` LIEF has been presented at [RMLL](https://prog2017.rmll.info/programme/securite-entre-transparence-et-opacite/lief-bibliotheque-d-instrumentation-de-formats-executables-mais-ca-fait-bife-c?lang=en) and the [MISP](http://www.misp-project.org) project uses it for its *PyMISP objects*. Some may complain about the C API. They are right! Until the ``v1.0.0`` we will provide a minimal C API. Once C++ API is stable we plan to provide full APIs for Python, C, Java, OCaml [^6], etc. Next version should be focused on the Mach-O builder especially for adding sections and segments. We also plan to support PE ``.NET`` headers and fix some performances issues. For questions you can join the [Gitter channel](https://gitter.im/lief-project) [^6]: https://github.com/aziem/LIEF-ocaml --- # LIEF - Library to Instrument Executable Formats > How LIEF was open-sourced to provide a cross-platform API for parsing, inspecting, and modifying ELF, PE, and Mach-O executable formats. - Canonical URL: https://lief.re/blog/2017-04-18-lief/ - Markdown: https://lief.re/blog/2017-04-18-lief/index.md - Authors: Romain Thomas - Published: 2017-04-18T00:00:00Z - Modified: 2017-04-18T00:00:00Z - Tags: ELF, PE, Mach-O, architecture **Tl;DR** LIEF is a library to parse and manipulate ELF, PE and Mach-O formats. Source code is available on [GitHub](https://github.com/lief-project/LIEF) and use cases are [here](http://lief.quarkslab.com/doc/latest/tutorials/index.html). ## Executable File Formats in a Nutshell When dealing with executable files, the first layer of information is the format in which the code is wrapped. We can see an executable file format as an envelope. It contains information so that the postman (i.e. Operating System) can handle and deliver (i.e. execute) it. The message wrapped by this envelope would be the machine code. There are mainly three mainstream formats, one per OS: * **Portable Executable (PE)** for Windows systems * **Executable and Linkable Format (ELF)** for UN*X systems (Linux, Android...). * Mach-O for OS-X, iOS... Other executable file formats, such as ``COFF``, exist but they are less relevant. Usually each format has a header which describes at least the target architecture, the program's entry point and the type of the wrapped object (executable, library...) Then we have blocks of data that will be mapped by the OS's loader. These blocks of data could hold machine code (``.text``), read-only data (``.rodata``) or other OS specific information. For PE there is only one kind of such block: **Section**. For ELF and Mach-O formats, a section has a different meaning. In these formats, sections are used by the **linker** at the **compilation** step, whereas **segments** (second type of block) are used by the OS's loader at **execution** step. Thus, sections are not mandatory for ELF and Mach-O formats and can be removed without affecting the execution. ### Purpose of LIEF It turns out that many projects need to parse executable file formats but don't use a *standard* library and re-implement their own parser (and the wheel). Moreover, these parsers are usually bound to one language. On Unix system one can find the ``objdump`` and ``objcopy`` utilities but they are limited to Unix and the API is not user-friendly. The purpose of LIEF is to fill this void: * Providing a cross-platform library which can parse and modify (to a certain extent) ELF, PE, and Mach-O formats using a common abstraction * Providing an API for different languages (Python, C++, C, and others) * Abstract common features from the different formats (Section, header, entry point, symbols...) The following snippets show how to obtain information about an executable using different API of LIEF: ```python import lief # ELF binary = lief.parse("/usr/bin/ls") print(binary) # PE binary = lief.parse("C:\\Windows\\explorer.exe") print(binary) # Mach-O binary = lief.parse("/usr/bin/ls") print(binary) ``` With the ``C++`` API: ```cpp #include int main(int argc, const char** argv) { LIEF::ELF::Binary* elf = LIEF::ELF::Parser::parse("/usr/bin/ls"); LIEF::PE::Binary* pe = LIEF::PE::Parser::parse("C:\\Windows\\explorer.exe"); LIEF::MachO::Binary* macho = LIEF::MachO::Parser::parse("/usr/bin/ls"); std::cout << *elf << std::endl; std::cout << *pe << std::endl; std::cout << *macho << std::endl; delete elf; delete pe; delete macho; } ``` And finally with the ``C`` API: ```c #include int main(int argc, const char** argv) { Elf_Binary_t* elf_binary = elf_parse("/usr/bin/ls"); Pe_Binary_t* pe_binary = pe_parse("C:\\Windows\\explorer.exe"); Macho_Binary_t** macho_binaries = macho_parse("/usr/bin/ls"); Pe_Section_t** pe_sections = pe_binary->sections; Elf_Section_t** elf_sections = elf_binary->sections; Macho_Section_t** macho_sections = macho_binaries[0]->sections; for (size_t i = 0; pe_sections[i] != NULL; ++i) { printf("%s\n", pe_sections[i]->name) } for (size_t i = 0; elf_sections[i] != NULL; ++i) { printf("%s\n", elf_sections[i]->name) } for (size_t i = 0; macho_sections[i] != NULL; ++i) { printf("%s\n", macho_sections[i]->name) } elf_binary_destroy(elf_binary); pe_binary_destroy(pe_binary); macho_binaries_destroy(macho_binaries); } ``` LIEF supports FAT-MachO and one can iterate over binaries as follows: ```python import lief binaries = lief.MachO.parse("/usr/lib/libc++abi.dylib") for binary in binaries: print(binary) ``` **Note** The above script uses the `lief.MachO.parse` function instead of the `lief.parse` function because `lief.parse` returns a **single** `lief.MachO.binary` object whereas `lief.MachO.parse` returns a **list** of `lief.MachO.binary` (according to the FAT-MachO format). Along with standard format components like headers, sections, import table, load commands, symbols, etc. LIEF is also able to parse PE Authenticode: ```python import lief driver = lief.parse("driver.sys") for crt in driver.signature.certificates: print(crt) ``` ```bash Version: 3 Serial Number: 61:07:02:dc:00:00:00:00:00:0b Signature Algorithm: SHA1_WITH_RSA_ENCRYPTION Valid from: 2005-9-15 21:55:41 Valid to: 2016-3-15 22:5:41 Issuer: DC=com, DC=microsoft, CN=Microsoft Root Certificate Authority Subject: C=US, ST=Washington, L=Redmond, O=Microsoft Corporation, CN=Microsoft Windows Verification PCA ... ``` Full API documentation is available here * [Python API](http://lief.quarkslab.com/doc/latest/api/python/index.html) * [C++ API](http://lief.quarkslab.com/doc/latest/api/cpp/index.html) * [C API](http://lief.quarkslab.com/doc/latest/api/c/index.html) ## Architecture In the ``LIEF`` architecture, each format implements at least the following classes: * **Parser**: Parse the format and decompose it into a ``Binary`` class * **Binary**: Modelize the format and provide an API to modify and explore it. * **Builder**: Transform the binary object into a valid file. ![Architecture](https://lief.re/blog/2017-04-18-lief/archi.png) To factor common characteristics in formats we have an inheritance relationship between these characteristics. For symbols it gives the following diagram: ![LIEF Symbol Inheritance](https://lief.re/blog/2017-04-18-lief/symbol_inheritance.png) It enables to write cross-format utility like ``nm``. ``nm`` is a Unix utility to list symbols in an executable. The source code is available here: [binutils](https://github.com/gittup/binutils/blob/0af702d47a443acea853b84157c2e81f6c131e77/binutils/nm.c) With the given inheritance relationship one can write this utility for the three formats in a single script: ```python import lief import sys def nm(binary): for symbol in binary.symbols: print(symbol) return 0 if __name__ == "__main__": r = nm(sys.argv[1]) sys.exit(r) ``` ## Conclusion As LIEF is still a young project we hope to have feedback, ideas, suggestions, and pull requests. The source code is available here: https://github.com/lief-project (under Apache 2.0 license) and the associated website: http://lief.quarkslab.com If you are interested in use cases, you can take a look at these tutorials: * [Parse and manipulate formats](http://lief.quarkslab.com/doc/latest/tutorials/01_play_with_formats.html) * [Create a PE from scratch](http://lief.quarkslab.com/doc/latest/tutorials/02_pe_from_scratch.html) * [Play with ELF symbols](http://lief.quarkslab.com/doc/latest/tutorials/03_elf_change_symbols.html) * [Hooking](http://lief.quarkslab.com/doc/latest/tutorials/04_elf_hooking.html) * [Infecting the PLT/GOT](http://lief.quarkslab.com/doc/latest/tutorials/05_elf_infect_plt_got.html) The project will be presented at the [Third French Japanese Meeting on Cybersecurity](http://cyber.science-japon.org/) ### Contact * lief [at] quarkslab [dot] com * Gitter: [lief-project](https://gitter.im/lief-project) ### Thanks Thanks to Serge Guelton and Adrien Guinet for their advice about the design and their code review. Thanks to Quarkslab for making this project open-source.