Stack-based buffer overflows: a practical introduction
Most exploit dev guides jump straight to writing a payload. This one stays on the mental model first — once the stack layout actually makes sense, the payload-writing part stops being a recipe you memorise and starts being something you can derive from first principles for any binary.
What’s on the stack during a function call
When a function is called, several things end up on the stack in a predictable order (x86-64, standard calling convention, no optimisations getting in the way):
higher addresses
┌─────────────────────────┐
│ caller's stack frame │
├─────────────────────────┤
│ saved return address │ ← where execution resumes after this function returns
├─────────────────────────┤
│ saved base pointer (RBP) │ ← caller's frame pointer, restored on return
├─────────────────────────┤
│ local variables │ ← buffers, counters, whatever the function declares
└─────────────────────────┘
lower addresses
The stack grows downward — toward lower addresses — as functions are called. Local variables sit at the bottom of this layout, closest to where a buffer write would start, and the saved return address sits above them, closer to where the write would end up if it goes far enough.
This ordering is the entire reason buffer overflows are dangerous: writing past the end of a local buffer doesn’t crash immediately, it keeps writing upward through memory that was never meant to receive untrusted data — including the address the CPU will jump to when this function returns.
Why a buffer overflow becomes a control-flow problem
A function like:
void greet() {
char name[64];
gets(name); // no bounds checking — reads until newline, however long that is
printf("Hello, %s\n", name);
}
gets() doesn’t know name is 64 bytes. It writes however much input it’s given. Input longer than 64
bytes keeps writing past the buffer — into the saved base pointer, then into the saved return address.
When greet() finishes and executes ret, the CPU doesn’t re-derive where to go. It trusts whatever 8
bytes are sitting at the saved return address slot. If that slot now contains an address you chose, the CPU
jumps there. That’s the entire exploit, conceptually — everything else is just figuring out which address
to put there and how many bytes of padding get you to that exact slot.
Finding the exact offset
Don’t guess with repeated As — use a cyclic (de Bruijn) pattern, where every 4 or 8-byte window in the
pattern is unique:
from pwn import cyclic, cyclic_find
pattern = cyclic(200)
# feed `pattern` as input, let the program crash
# read whatever ended up in RIP/EIP at the crash
offset = cyclic_find(0x6161616c) # the value that was in RIP
cyclic_find reverses the pattern back into a byte offset. That offset is the exact distance from the start
of your input to the saved return address — every payload after this point starts with that many bytes of
padding before the part that matters.
What replaces the address depends on what’s available
Once you control the saved return address, what you put there depends on the binary’s protections and what’s reachable:
- No NX (executable stack) — point to shellcode you placed on the stack yourself
- NX enabled, no ASLR (or a leak available) — build a ROP chain from gadgets already in the binary or
libc, or jump straight to an existing function like
system()if its address is fixed and known - Full protections (NX + ASLR + canary) — need an information leak before any of the above is possible; this is the step where most modern exploit chains spend the bulk of their effort
Where this guide stops
This is the conceptual layer — what’s true regardless of which specific binary you’re looking at. The practical version of “build a ROP chain” or “leak a libc address” depends entirely on the protections in front of you, which is why those are separate, narrower guides rather than sections bolted onto this one. The stack-smash-101 lab is the natural next step — it’s exactly this model applied to a real binary with no protections at all.