Level 0 · Chapter 2
The PE format
Headers, sections, imports, relocations — the on-disk contract every Windows loader decision is based on.
The loader has to make a lot of decisions — where to map an image, which DLLs to pull in, whether relocations are needed. It doesn't guess at any of them. Every decision is driven by data encoded directly in the file, in a format called PE (Portable Executable). Understanding PE is understanding the contract between whoever compiled and linked your program and the loader that will eventually run it.
The shape of a PE file
Every .exe, .dll, and .sys on Windows is a PE file, and they all share the same skeleton:
- The MS-DOS stub. The very first bytes are still a tiny MS-DOS program — historically, one that just printed "This program cannot be run in DOS mode." Windows ignores it functionally, but keeps it for backward compatibility; it's a small, permanent fossil of Windows' own history.
- The PE header. Immediately after the DOS stub, a signature (
PE\0\0) and the COFF header — machine type (x86-64, ARM64, ...), number of sections, timestamp, and characteristics flags (is this a DLL? is it valid for a given subsystem?). - The optional header (which, despite the name, is always present in practice). This is where the loader finds the entry point address, the preferred base address, the size of the image, the subsystem (GUI vs. console), and — critically — the data directories: fixed-size pointers to the import table, export table, resource table, relocation table, and several others.
- Section headers. A list describing each named region of the file — typically
.text(code),.data(initialized variables),.rdata(read-only data, including the import table),.rsrc(resources like icons and version info),.reloc(relocation data) — along with where each section should be mapped in memory and what protection (read/write/execute) it needs.
Why sections exist as their own concept
A naive loader could just map the whole file into memory as one blob and jump to an offset. PE doesn't do that, and the reason is security and correctness, not tradition: code and data need different memory protections. If .text and .data were mapped as one undifferentiated region, it would have to be both executable and writable — which is exactly the combination attackers want, because writable+executable memory is what lets a data corruption bug turn into arbitrary code execution. By splitting the image into sections with per-section permissions, Windows can map .text as read+execute only, so even a successful buffer overflow into code can't just write new instructions there.
Imports: the table the loader actually reads
The import table is a per-DLL list of function names (or ordinals) a module depends on. For each entry, the linker leaves a slot — the Import Address Table (IAT) — that the compiled code calls through indirectly. At load time, the loader:
- Finds each imported DLL (via the search order covered in the DLL loader chapter).
- Looks up each imported function's real address in that DLL's export table.
- Writes that address into the IAT slot.
This indirection is why a program compiled against one version of kernel32.dll still works against a newer one with a different internal layout: the code never hardcodes an address, only an IAT slot that gets filled in fresh on every load.
Relocations: what happens when the preferred address is taken
Every PE image has a preferred base address baked in at link time (historically 0x00400000 for EXEs). If nothing else occupies that address when the image loads, great — no further work needed. But with ASLR enabled (the default since Windows Vista), Windows deliberately picks a randomized base address for security, specifically to make return-oriented-programming exploits harder to write reliably.
When an image loads somewhere other than its preferred address, every absolute address the compiler baked into the code — a pointer to a global variable, for instance — is now wrong by a fixed offset. The .reloc section is a list of exactly which locations in the image need that offset added. The loader walks this list and patches each one. This is pure overhead that exists solely to make ASLR possible, which is a fair trade: predictable load addresses are a gift to exploit authors.
A common mistake
It's tempting to think of a PE file as "the code, plus some metadata Windows ignores." In practice the metadata is the interface — the loader's entire behavior (where to map things, what to load, how to patch addresses) is driven by data directories, section flags, and table entries, not by inspecting the actual instructions. Tools like dumpbin /headers or CFF Explorer are popular specifically because they show you the same data the loader itself is reading.
Where this connects
- Every DLL your program depends on is discovered and mapped by the DLL loader, which reads exactly the import-table structure described here.
- The sections this chapter describes are placed into the process address space by Memory management — specifically tracked as regions in the VAD tree.