Systems · Concurrency · Deep Dive

A journey from the kernel's process table to the event loop

Written for the endlessly curious student — the full story of how computers wait, and why every language tells the same lie at a different layer.

14 sections ~25 min read Plain English + code
01

🏗️ The Process — The Container of Everything

Every program that runs on your computer is a process. A process is not your code — it's the container the OS creates to hold:

  • 📄 Your code (in memory)
  • 🧮 A stack (local variables, return addresses)
  • 📦 A heap (dynamically allocated memory)
  • 🔌 Open file descriptors (sockets, files, pipes)
  • 🏷️ A PID (Process ID)
  • 📊 A state (running, sleeping, stopped…)

🏷️ The PID

A PID is a unique integer the kernel assigns to every process. Key facts:

  • PID 1 is always init or systemd — the ancestor of all processes.
  • PIDs are assigned sequentially and recycled after reaching pid_max (typically 32,768 or 4,194,304).
  • Every process has a PPID (Parent PID). If a parent dies, children are reparented to PID 1.
  • Your shell has a PID. Type echo $$ to see it.

A PID is not "the process" — it's just a handle the kernel uses to refer to it.

🔄 Process States

At any instant, a process is in exactly one of these states:

StateSymbolMeaningCPU Usage
RunningRExecuting or ready to executeActively using CPU
SleepingSWaiting for I/O or event0%
Uninterruptible SleepDWaiting for disk I/O0%
ZombieZFinished, but parent hasn't reaped it0%
StoppedTPaused (e.g., Ctrl+Z)0%

You can see these in ps aux or top. A healthy server spends 99.9% of its time in S — sleeping, waiting, doing nothing. That's not a bug. That's the entire design.

02

♾️ Infinite Loops — Why Servers Never Die

Here's a truth that surprises many beginners: a process ends when its code ends. When your script reaches the last line, the OS frees its RAM, closes its file descriptors, and removes its PID. Done.

So how does a server stay alive? With an infinite loop:

while (true) {
    $data = fread($socket, 1024);
    process($data);
}

But wait — doesn't that burn CPU? This is where the crucial distinction lives:

🚫 Bad Infinite Loop (Busy-Waiting)

while (true) {
    $x = 1 + 1; // doing nothing useful
}

This runs 100% CPU. The OS scheduler sees it as "runnable" and keeps giving it time slices. This does hurt the system.

✅ Good Infinite Loop (Blocking)

while (true) {
    $data = fread($socket, 1024); // BLOCKS until data arrives
    process($data);
}

When fread() finds no data, the kernel puts the process to sleep. The process enters state S and is removed from the CPU run queue entirely. It uses 0% CPU. When data arrives, the kernel wakes it up.

Key insight

A well-written infinite loop spends 99.9% of its time sleeping inside a blocking syscall, not spinning. The loop exists only to re-enter the blocking call.

Every long-running service on Earth uses this pattern:

ServiceWhere it blocks
Nginxepoll_wait()
Node.jsepoll_wait() (via libuv)
Swooleepoll_wait() (Reactor threads)
sshdaccept()
Your PHP scriptfread() on a socket
systemdepoll_wait()

The pattern is always the same: block → wake on work → process → block again.

03

🌊 The Fundamental Truth — I/O Always Blocks

Here is the single most important fact in this entire article:

The fundamental truth Blocking is not a language flaw. It is a hardware reality.

⏱️ Why Blocking Is Unavoidable

A CPU executes billions of instructions per second. But I/O devices operate on a completely different timescale:

OperationLatencyCPU cycles (at 3GHz)
L1 cache access~1 ns3
RAM access~100 ns300
SSD read~100 µs300,000
Network round-trip (LAN)~500 µs1,500,000
Network round-trip (internet)~50 ms150,000,000
Disk seek (HDD)~10 ms30,000,000

When your code says "read from this socket," the data is not there yet. The CPU has exactly two choices:

  1. Spin — execute a loop checking "is it here yet?" → burns 100% CPU for millions of cycles doing nothing.
  2. Sleep — tell the kernel "wake me when data arrives" → 0% CPU, but the process is paused.

Blocking is the CPU's way of not wasting itself on waiting. It's the efficient answer to an unavoidable physical gap.

🎭 The Two Choices Visualized

❌ Busy-wait (spin):

CPU: is it ready? is it ready? is it ready? ... [1,000,000 times] ... yes!
CPU usage: 100% 😱

✅ Blocking (sleep):

CPU: kernel, wake me when ready. [SLEEPS] ... kernel: WAKE UP!
CPU usage: 0% 😌

There is no third option. The data physically cannot arrive faster than the network allows.

04

📡 Signals — How the OS Interrupts You

A signal is an asynchronous notification sent to a process by the kernel or another process. It's the OS's way of saying "hey, something happened."

⌨ How Ctrl+C Works

When you press Ctrl+C in a terminal:

  1. The terminal driver sends SIGINT (signal 2) to the foreground process group.
  2. The process either:
    • Dies immediately (default behavior), or
    • Catches the signal and runs a handler (graceful shutdown).

This is why even an infinite loop dies when you press Ctrl+C — the OS interrupts it. Your loop was sleeping in fread(), and the signal wakes it with a special error.

📋 Common Signals

SignalNumberDefault ActionTriggered By
SIGINT2TerminateCtrl+C
SIGTERM15Terminate (graceful)kill PID (default)
SIGKILL9Terminate immediately (cannot be caught)kill -9 PID
SIGHUP1TerminateTerminal closes
SIGSTOP19Pause (cannot be caught)kill -STOP PID
SIGCONT18Resumekill -CONT PID
SIGUSR110User-definedkill -USR1 PID

🎯 The Special Case: SIGKILL

Cannot be caught

SIGKILL cannot be caught, blocked, or ignored. The kernel terminates the process immediately — no cleanup, no file flushing, no graceful goodbyes. This is why you use it as a last resort.

🛠️ Killing Processes on Linux

ps aux              # list all processes
top                 # live view, sorted by CPU
pgrep -a php        # find processes named "php"
pstree -p           # tree view with PIDs
kill 1234           # graceful terminate (SIGTERM)
kill -9 1234        # force kill (SIGKILL)
pkill php           # kill all named "php"
killall php         # same, by exact name
05

🧵 Threads — The OS's Original Concurrency Tool

A thread is the smallest unit of execution the kernel can schedule. It consists of:

  • A program counter (where in the code it is)
  • A stack (local variables, return addresses) — typically 1–8 MB
  • A set of CPU registers
  • An entry in the kernel's scheduler

🔄 Context Switching

When the OS switches threads, it must:

  1. Save all registers of the current thread
  2. Save the stack pointer
  3. Load registers of the next thread
  4. Load its stack pointer
  5. Flush TLB entries

This costs 1–10 microseconds. Sound fast? A CPU executes ~3,000–30,000 instructions in that time. It's expensive.

⚡ Multithreading

Multiple threads inside one process, sharing the same memory space:

Process
├── Thread 1  ──┐
├── Thread 2  ──├ shared heap, shared globals
├── Thread 3  ──├ separate stacks
└── Thread 4  ──┘

Pros:

  • ✅ True parallelism on multiple cores
  • ✅ Shared memory = fast data exchange

Cons:

  • ❌ Race conditions — two threads writing the same variable
  • ❌ Deadlocks — two threads each waiting for the other's lock
  • ❌ Memory cost — 8 MB × 1000 threads = 8 GB just for stacks
  • ❌ Debugging is hard — bugs are non-deterministic
  • ❌ Limited scale — a few thousand threads per machine before things degrade

⚠️ Why This Doesn't Scale

The wall

Want 100,000 WebSocket connections with one thread each? That's 800 GB of stack memory. Impossible. This is why the industry moved to lighter-weight abstractions.

06

🚪 The Kernel's Waiting Room — epoll, kqueue, IOCP

The Problem

Imagine 10,000 open sockets. You want to know which one has data. The naive approach:

// BAD: old select() approach
for (int i = 0; i < 10000; i++) {
    if (has_data(sockets[i])) {  // check each one
        read(sockets[i]);
    }
}

That's O(n) per check — 10,000 syscalls just to find one ready socket. Terrible.

The Solution: epoll (Linux), kqueue (BSD/macOS), IOCP (Windows)

These are kernel mechanisms that let one thread watch many file descriptors at once:

int epfd = epoll_create1(0);
epoll_ctl(epfd, EPOLL_CTL_ADD, socket1, &event);  // register once
epoll_ctl(epfd, EPOLL_CTL_ADD, socket2, &event);
// ... register all 10,000

struct epoll_event events[100];
int n = epoll_wait(epfd, events, 100, -1);  // BLOCKS here

// returns ONLY the ready sockets, e.g. 3 of them
Key insight

epoll_wait blocks (thread sleeps at 0% CPU) until at least one registered fd has data. Then it returns only the ready ones. You registered 10,000 sockets, you get back 3.

🌍 Platform Equivalents

PlatformMechanismSince
Linuxepoll2002
macOS / BSDkqueue2000
WindowsIOCP1994
Solarisevent ports2000
Cross-platform (older)select / poll1980s

This is a kernel feature. No language "has" epoll — every language's runtime uses it on Linux.

07

🤝 Coroutines & Goroutines — Cooperative Multitasking

A coroutine is a function that can pause itself and be resumed later. That's the whole idea.

🤝 Threads vs Coroutines

ThreadCoroutine
Who pauses itOS (preemptive)Itself (cooperative)
When it pausesAny instructionOnly at yield/await points
Stack size1–8 MBA few KB (or zero)
Managed byKernelLanguage runtime
Context switch~1–10 µs~10–100 ns
How manyThousandsMillions
Data racesYesOnly if multi-threaded

Cooperative scheduling is the key. A coroutine runs until it explicitly says "I'm waiting for something." No preemption means no data races between coroutines on the same thread.

📚 Two Flavors

  • Stackful — each coroutine has its own tiny stack. Go, Lua, Swoole.
  • Stackless — compiled into a state machine; no separate stack. JS async/await, Rust async, C++20.

🐹 Goroutines — Go's Special Thing

A goroutine is Go's stackful coroutine with M:N scheduling:

Goroutines:  G1 G2 G3 G4 G5 G6 G7 G8 ... G1000000
              \  |  /    \  |  /
OS threads:    T1        T2        T3        T4
                \        |         /
CPU cores:      Core 1  Core 2   Core 3   Core 4

When a goroutine blocks on I/O, the Go runtime parks it and runs another goroutine on the same OS thread. The OS thread doesn't block — the runtime handles waiting via epoll/kqueue/IOCP.

The programming model is beautiful: you write code that looks blocking:

func handle(conn net.Conn) {
    buf := make([]byte, 1024)
    n, _ := conn.Read(buf)   // looks blocking
    process(buf[:n])
}

for {
    conn, _ := listener.Accept()
    go handle(conn)          // spawn a goroutine per connection
}

No async, no await, no callbacks. You can have millions of goroutines, each costing ~2 KB.

08

🎁 Promises & Async/Await — Sugar Over State Machines

🧩 What a Promise Actually Is

A Promise is not a mechanism. It's a value — an object with three states and a list of callbacks:

┌─────────────────────────────────────────┐
│  Promise                                 │
│  state: "pending"                        │
│         "fulfilled"  (with a value)      │
│         "rejected"   (with a reason)     │
│  callbacks: [fn1, fn2, fn3, ...]         │
└─────────────────────────────────────────┘

It does no I/O. It doesn't know about sockets. It's just a container waiting for someone to call resolve or reject.

const promise = new Promise((resolve, reject) => {
    const socket = connect(url);
    socket.onData = (data) => resolve(data);   // someone calls resolve
    socket.onError = (err) => reject(err);
});

The Promise sits in the gap while epoll_wait is blocking. When data arrives, the event loop calls resolve. The state changes to fulfilled, and all stored callbacks fire.

✨ await — Just Sugar

await is not a primitive. It's a keyword that the compiler turns into a state machine:

// You write:
async function getData() {
    const res = await fetch(url);
    const data = await res.json();
    return data;
}

The compiler transforms this into something like:

function getData() {
    return new Promise((resolve) => {
        let state = 0;
        function step(value) {
            switch (state) {
                case 0: state = 1; fetch(url).then(step); return;
                case 1: state = 2; value.json().then(step); return;
                case 2: resolve(value); return;
            }
        }
        step();
    });
}

Key facts about await:

  • The function stops executing at await.
  • Its locals are saved in a heap-allocated state machine.
  • Control returns to the event loop, which runs other tasks.
  • When the awaited Promise resolves, the runtime resumes the function.

🎭 Promise vs await

Promiseawait
What it isA value (object)A keyword (syntax)
When it existsAt runtimeAt compile time
What it doesHolds a future resultPauses the function
Can you have one without the other?✅ Yes (Promises existed before await)❌ No (await needs a Promise)

await is sugar. Underneath, it's the same Promise mechanism.

⚡ Eager vs Lazy

  • JS Promise: eager. new Promise(() => doWork()) runs doWork() immediately.
  • Rust Future: lazy. let f = async { do_work().await }; does nothing until you .await.

This is why Rust's async feels so different — the Future is a recipe, not a running task.

09

🎩 libuv — The Hidden Loop Inside Node.js

libuv is the C library that powers Node.js's event loop. It's the "hidden loop" we discussed earlier — the thing that makes ws.onmessage feel effortless.

🎯 What It Abstracts

libuv picks the best I/O polling backend for each platform:

  • Linux → epoll
  • macOS / BSD → kqueue
  • Windows → IOCP

🧠 The Truth About "Non-Blocking"

When people say "Node.js is non-blocking," they mean something specific: the JavaScript thread never blocks. But something is still blocking:

Your JS code:      ws.onmessage = handler
                   ↓
Node's event loop: epoll_wait(...)   ← BLOCKS HERE, 0% CPU
                   ↓
Kernel:            wakes when socket has data
                   ↓
libuv:             reads data, parses frame
                   ↓
Your handler:      runs on the JS thread
Remember

The epoll_wait() call blocks. The OS puts the Node process to sleep. "Non-blocking" is always a lie told at a higher level of abstraction.

10

🌐 WebSockets in Practice — JS vs PHP vs Swoole

📜 The Truth: Every WebSocket Client Has a Loop

There is no magic. To receive data continuously, something must repeatedly check "is there new data?" That something is always a loop.

🟥 JavaScript — The Loop Is Hidden

const ws = new WebSocket('wss://stream.binance.com:9443/ws/btcusdt@trade');

ws.onmessage = (event) => console.log(event.data);

You never see a loop. But underneath, V8 + libuv is doing something like:

while (true) {
    events = epoll_wait(...);   // blocks
    for (event : events) {
        data = read(event.socket);
        frames = parse_websocket_frames(data);
        for (frame : frames) {
            dispatch_to_js_callback(frame);  // calls your onmessage
        }
    }
}

🐘 PHP (Raw Sockets) — You Write the Loop

while (!feof($socket)) {
    $header = fread($socket, 2);   // blocks
    $payload = fread($socket, $len);
    echo decodeFrame($header . $payload)['payload'] . "\n";
}

This is the same loop that JS hides. You just wrote it manually.

🚀 PHP (Swoole) — The Loop Comes Back, Coroutine-Style

Co\run(function() {
    $client = new Co\Http\Client('stream.binance.com', 9443, true);
    $client->upgrade('/ws/btcusdt@trade');

    while (true) {
        $frame = $client->recv();  // coroutine-yields, doesn't block the process
        echo $frame->data . "\n";
    }
});

Swoole's coroutine runtime uses epoll internally, pausing the coroutine while waiting and running others.

📊 The Three Compared

ConceptJavaScriptPHP (raw)PHP (Swoole)
The loopHidden in V8You write whileRuntime handles it
Receiving dataws.onmessage = fnfread()$client->recv()
Blocking?Non-blockingBlockingNon-blocking
Who pollslibuvYour loopSwoole Reactor
Concurrent connectionsThousandsOne per processThousands per process
11

🏎️ Frameworks & Runtimes — Octane, Swoole, OpenSwoole

🐘 Laravel Octane — A Performance Multiplier

Octane boots your Laravel app once and keeps it in memory, then feeds it requests at high speed. It's a request-serving layer, not a non-blocking runtime.

The catch: Octane is only as non-blocking as the server you pair it with.

Octane ServerCoroutine SupportBlocking I/O?
Swoole (standard)Partial⚠️ Still blocks per worker
Swoole + Coroutine Hooks✅ Yes✅ True non-blocking
RoadRunner❌ No⚠️ Blocking
FrankenPHP❌ No⚠️ Blocking

Without coroutine hooks, Octane with Swoole is still blocking — 8 workers with 1s I/O = 8 req/s. With coroutine hooks, those same 8 workers can handle thousands of concurrent requests.

🚀 Swoole vs OpenSwoole

Swoole is the original. OpenSwoole is a community fork. Both are C extensions providing:

  • Coroutines (stackful, like Go)
  • Multi-process + multi-threaded Reactor model
  • Coroutine-aware I/O (Co\MySQL, Co\Http\Client, etc.)
  • Runtime hooks that convert standard PHP blocking functions into coroutine-aware versions

The magic hooks:

OpenSwoole\Runtime::enableCoroutine(OPENSWOOLE_HOOK_ALL);

Co\run(function() {
    sleep(5);  // ← converted to non-blocking coroutine sleep!
    // Process is NOT blocked. Other coroutines keep running.
});

This works for sleep(), file_get_contents(), PDO, Redis, cURL, and more.

📊 Traditional PHP vs Swoole — 3 Tasks, 1 Second Each

ModelProcessesWorkersTotal Time
Traditional PHP (blocking)33~3 seconds
Swoole (coroutines)11~1 second
12

🌍 The Grand Comparison — Every Language Side by Side

PHP (FPM) PHP (Swoole) Go JS (Node) C (pthreads) Java (threads) Java 21 (virtual) Rust (tokio)
Unit ProcessCoroutineGoroutineAsync task ThreadThreadVirtual threadAsync task
Scheduler OSSwooleGo runtimelibuv OSOSJVMtokio
Preemptive? ✅❌❌❌ ✅✅❌❌
Stackful? N/A✅✅❌ ✅✅✅❌
Stack size ~8 MB~2 KB~2 KBheap ~8 MB~1 MB~few KBheap
Max count ~thousands~millions~millions~millions ~thousands~thousands~millions~millions
True parallel? ✅✅✅❌ ✅✅✅✅
Shared memory? ❌✅✅❌ ✅✅✅✅*
Data races? ❌PossiblePossible❌ ✅✅Possible✅*
Runtime provided? N/A✅✅✅ ❌✅✅❌

* Rust uses Send/Sync to guarantee no data races at compile time.

💰 Cost Numbers That Matter

UnitCreation timeMemory
OS thread~10–100 µs1–8 MB
Goroutine~200 ns2 KB
JS async task~100 ns~few hundred bytes
Rust async task~50 ns~few hundred bytes
Swoole coroutine~1 µs~few KB
Switch typeTime
OS thread → OS thread1–10 µs
Goroutine → Goroutine100–200 ns
Async task → Async task50–100 ns

The industry trend is clear: move away from OS threads toward runtime-managed coroutines.

13

🔮 The Unifying Truth — Blocking Is Everywhere

Here is the mental model that ties everything together:

┌──────────────────────────────────────────────┐
│  Your code                                   │
│  await fetch(url)   yield value              │  ← language syntax
├──────────────────────────────────────────────┤
│  Language runtime                            │
│  Promises, state machines, schedulers        │  ← runtime
├──────────────────────────────────────────────┤
│  I/O library (libuv, tokio, Swoole, asyncio) │  ← library
├──────────────────────────────────────────────┤
│  Kernel (epoll, kqueue, IOCP)                │  ← kernel
│  epoll_wait() — BLOCKS a thread              │
├──────────────────────────────────────────────┤
│  Hardware                                    │  ← physics
└──────────────────────────────────────────────┘

Every layer hides the blocking of the layer below it.

  • epoll_wait blocks the thread — but only one thread.
  • The event loop doesn't block your code — it runs other tasks while epoll_wait sleeps.
  • await doesn't block the event loop — it suspends your function and returns control.
  • yield doesn't block anything — it hands control back to the caller or scheduler.

The blocking never disappears. It just moves down the stack. Every "non-blocking" language is blocking somewhere — in a runtime, in a thread pool, or in the kernel.

🎯 The One-Sentence Answer

Every language blocks. The only difference is who writes the sleep and how many sleepers you can afford.
14

🎓 Conclusion — What to Take Away

If you remember nothing else from this article, remember these ten truths:

  1. A process is a container — code, memory, file descriptors, and a PID.
  2. Processes die when their code ends — servers stay alive with infinite loops.
  3. Good loops block, bad loops spin — one sleeps at 0% CPU, the other burns 100%.
  4. I/O blocking is physics, not a language flaw — the CPU is orders of magnitude faster than I/O devices.
  5. Signals are how the OS interrupts — Ctrl+C sends SIGINT; kill -9 sends SIGKILL.
  6. Threads are preemptive and expensive — a few thousand max, 1–8 MB each.
  7. Coroutines are cooperative and cheap — millions possible, a few KB each.
  8. epoll/kqueue/IOCP are kernel primitives — one thread watches many fds, blocks at 0% CPU.
  9. Promises are values, await is sugar — both rely on the event loop and epoll underneath.
  10. Non-blocking is always a lie at some layer — blocking gets moved, never eliminated.

🧭 The Final Mental Image

Imagine a restaurant:

  • epoll = the kitchen's order bell system (one panel, lights up when any table's order is ready).
  • Event loop = the waiter, standing by the panel, delivering orders as they light up.
  • Promise = the ticket the kitchen hands back — "your food will be ready."
  • await = a customer saying "I'll order dessert after the main" — the waiter serves others, comes back later.
  • Coroutine = a chef who pauses plating one dish to start another, resuming seamlessly.

The bell is in the kitchen. The waiter is in the dining room. The customer's request is at the table. Three layers, all connected, none the same thing.

The world of concurrency is not about eliminating waiting. It's about waiting efficiently — and every language, framework, and runtime is just a different answer to the same fundamental question:

"While one thing waits, what can the rest do?"