In theory, a generation upgrade in Apple Silicon sounds straightforward enough to outsiders: new silicon, refined process node, slightly higher IPC, perhaps a few extra execution units. Underneath, it remains 64-bit ARM. Anyone running Linux on standard x86 iron or generic server boards generally expects a working kernel to survive at least until early console output on the next processor iteration.
That assumption does not hold on Apple's platform. There is neither BIOS nor ACPI—no standardized firmware layer that hands the OS a map of memory, timers, or serial ports at boot. Every new chip generation starts as a black box, forcing you to grope in the dark for the switches.
Just how tedious and granular that process is in practice is illustrated by a technical write-up from developer Yureka Lilian: The forgetful CPU (Linux on M4). In it, she documents the uphill battle of coaxing Linux into dropping to a bare shell on an M4 Mac mini. It offers an unvarnished look at the reality of low-level hardware bring-up: the job does not start with tidy documentation, but with raw low-level debugging, dead-silent serial consoles, and processors that quietly wipe their own registers when idling.
The m1n1 Foundation and New Barriers
To understand why a kernel fires up on a Mac in the first place, you have to look at the groundwork: m1n1. It serves as both the bootloader and experimental testbed for the Asahi Linux project. Normally, m1n1 isn't merely the glue between Apple's bootchain (iBoot) and U-Boot or Linux; it also doubles as a lightweight hypervisor.
On earlier generations (M1 through M3), that hypervisor was the key to porting:
macOS was booted inside this virtualized wrapper while m1n1 recorded every MMIO access in the background—the raw reads and writes where the OS addresses hardware controllers like Display, PCIe, or USB directly across dedicated memory addresses (Memory-Mapped I/O). These execution traces revealed exactly how the undocumented silicon expected to be driven.
With the M4, that approach hit a new roadblock. Starting with this chip generation, Apple made the Secure Page Table Monitor (SPTM) mandatory—a security subsystem enforcing page table integrity inside Apple's XNU kernel. That effectively broke the established hypervisor tracing route. Getting macOS to boot under m1n1 for clean register tracing now demands far more intrusive changes to the virtualization layer.
That left the direct bare-metal route for initial bring-up:
Drop system security down to "Reduced Security" via macOS Recovery, register m1n1 as an alternate boot target, and attempt to drop Linux directly onto bare metal.
Right away, it became obvious how allergic the new platform is to routines that were standard operating procedure on older generations. Early m1n1 builds crashed instantly on the M4. As Yureka Lilian documents in her post, this came down to hardcoded register writes that Apple had locked down on the newer SoCs.
These included init sequences for Apple's Guarded Execution Framework (GXF) along with writes to RVBAR (Reset Vector Base Address Register), which governs where a CPU core jumps after power-on. Registers that required explicit setup on older chips were already locked down by hardware or hardwired to correct defaults on the M4. Attempting to overwrite them ended in an early crash before Linux had meaningfully entered the picture.
Only after m1n1 was patched to deliberately skip these write cycles did the bootloader return to life over the serial console.
A Single Character: Debugging in the Dark
If you've never brought up an operating system on raw, undocumented silicon, you might picture early boot failures as text on a screen—a flashing cursor, a kernel panic, or a cryptic error code. The reality of early bring-up is much bleaker:
Nothing happens at all.
Following the Vectoring to next stage command, where m1n1 hands execution off to the Linux kernel, the Mac mini's serial console stayed stone cold. No error code, no register dump, no sign of life.
Standard Linux troubleshooting tools are useless at this stage. When the kernel hasn't made it far enough to initialize drivers or stand up the standard console infrastructure, developers typically rely on earlycon. This boot parameter tells the kernel to route output to serial hardware registers (UART) as early as possible using minimal, hardcoded routines. But earlycon couldn't help either, as long as the kernel had no reliable way of knowing which UART controller it was supposed to talk to.
When you don't even know whether your code is executing or the CPU is hanging on cycle zero, you're left with the crudest tool available: assembly printf-debugging.
Yureka Lilian grabbed a bare-bones routine named debug_putc straight out of m1n1.
This isn't a complex driver; it's a handful of assembly instructions blasting a hardcoded byte directly to the physical memory address of the UART chip. In this case: a single ASCII character, the letter a.
She patched that routine into the very earliest entry phase of the ARM64 kernel (arch/arm64/kernel/head.S)—and against the odds, the serial console spat out an a.
It sounds trivial, but during bring-up, that's everything:
The CPU is executing actual Linux code. It is alive.
From there, bisection debugging began: moving the debug output line by line, instruction block by instruction block further down the boot path to pinpoint exactly where execution died. The trail quickly vanished at a notorious bottleneck: enabling the Memory Management Unit (MMU).
The MMU is the hardware block in the processor responsible for address translation. Before it's switched on, the CPU interacts directly with raw physical addresses in RAM and memory-mapped controllers. Once the MMU is active, the processor operates strictly in virtual address space, resolving every access against page tables to locate the actual physical backing.
The boot trap:
m1n1 had set up an identity mapping (1:1 translation) for its own environment, where virtual controller addresses matched physical memory locations. Linux does not do this out of the box in its early temporary page tables. The second the MMU engaged, debug_putc attempted to write to a virtual address that had no corresponding entry in the active tables. The write faulted into unmapped space—turning the emergency debug output into the very crash it was meant to trace.
Only after Lilian patched the kernel's initial page tables in source, hardcoding a 1:1 identity map for the serial controller's address range under active MMU, did that little a continue its journey down the boot pipeline.
Shortly after, the boot path stumbled over an implementation-specific Apple register write (SYS_IMP_APL_VM_TMR_FIQ_ENA_EL2) tied to early interrupt initialization. This virtualization-related register—unlocked by Apple in later iBoot firmware—triggered an immediate fault; only after commenting out the write access did the system boot through to a shell.
The mystery behind the dead earlycon output also unraveled cleanly:
By explicitly declaring the interface via earlycon=s5l,0x3ad200000 and setting stdout-path = "serial0" in the Device Tree—the hardware manifest informing the kernel about controllers, addresses, and interrupts—Linux finally started serving up what engineers actually need: full register dumps and stack traces on early panics.
A missing property in a hardware description, an unmapped address window, and days spent hunting down a single ASCII character on a terminal: That is what real hardware bring-up looks like on the front lines.
When WFI Doesn't Wait, It Wipes
Up to this point, Linux was running on a single core. But nobody buys modern hardware to use it as a uniprocessor system. As soon as you try to bring the remaining cores online and transition the OS into regular scheduling, the real architectural wall on the M4 presents itself. It also explains the title of Yureka Lilian's blog post:
The CPU simply forgets its architectural state.
In the ARM64 ISA, there is a fundamental instruction for this:
WFI (Wait For Interrupt), alongside newer extensions like WFIT (Wait For Interrupt with Timeout). The premise: when the kernel wants to place an idle CPU core into a low-power holding pattern, it routes execution through these WFI/WFIT paths. The core drops into a power-saving sleep until an external hardware event—a timer tick, device I/O, or an inter-processor interrupt—wakes it back up. Once triggered, execution resumes exactly where it stopped.
This relies on an axiomatic rule of computer architecture:
A power-down wait state must not destroy active working state. Lilian points to the ARM64 architecture reference manual in her blog post: when an implementation allows a WFI instruction to complete, the instruction must not cause a loss of architectural state. Whatever sat in the CPU registers before entering idle must remain intact upon wake-up.
On the M4, out of the box, that guarantee evaporates.
General-purpose registers x0 through x31—the working memory where the processor keeps pointers, addresses, loop iterators, and live arguments—are zeroed out when waking from sleep. For an operating system, this is catastrophic. The kernel sleeps for a few microseconds, wakes up, attempts to service an interrupt or execute the next task, finds zeros where live register state should be, and immediately panics. It doesn't merely crash a userland application; it cuts the ground out from under the running kernel mid-stride.
Strictly speaking, this quirk isn't entirely new to Apple Silicon.
Earlier generations (M1 through M3) exhibited a similar trait rooted in Apple's aggressive power gating: to drop cores into deep low-power states, the hardware scrubs state quickly. Apple's own operating system kernel, XNU, was engineered around this behavior from day one: it spills all relevant registers to memory before calling WFI, and restores them manually upon wake-up.
Furthermore, older chips featured a hardware bypass—an internal core override bit known as a chicken bit (ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask) that could suppress this aggressive power gating when debugging or running legacy code. The m1n1 bootloader previously toggled this bit on M1–M3 silicon, forcing the cores to behave like standard-compliant ARM64 CPUs out of the gate. The Asahi Linux kernel could then selectively manage power states later via targeted drivers.
On the M4, that chicken bit appears to be either locked or gone entirely. Apple builds hardware first and foremost for macOS and XNU, where software register saving has been production-hardened for years. Consequently, almost nobody outside Apple's ecosystem ever cared about this architectural idiosyncrasy. For Linux bring-up, however, the primary hardware knob used to enforce standard WFI semantics was completely off the table.
From a Quick Hack to a Clean Upstream Fix
The first working breakthrough in bring-up work is almost always crude. To verify whether WFI logic was indeed the sole roadblock preventing secondary cores from spinning up, Yureka Lilian reached for a blunt instrument: she patched the kernel to replace every single WFI and WFIT instruction with a NOP (No Operation).
A NOP does nothing. The processor doesn't wait, doesn't drop into sleep, preserves its registers, and simply steps to the next instruction.
The result: Linux booted cleanly across all CPU cores on the M4, straight to a shell.
As a proof of concept, this was priceless. As a permanent solution for a production system or Linux mainline, it was a non-starter: it burns continuous power, defeats proper idle states, and would turn laptop battery runtimes into a bad joke.
The next logical impulse would have been a traditional errata patch.
The ARM64 port of Linux has well-established mechanisms for patching around known silicon errata during early boot memory setup. But that ran straight into another trap:
How does the kernel definitively determine it is running on a bare-metal M4 that drops its registers?
On Macs, Linux doesn't just run bare-metal via m1n1; it frequently runs virtualized inside guests under macOS. The macOS hypervisor intercepts WFI executions from its VMs to yield CPU slices efficiently to other host processes. If the Linux kernel unconditionally neutered WFI based solely on CPU part numbers, it would penalize Linux guests running in macOS VMs. Detecting virtualization—especially nested virtualization—at early boot is notoriously fragile.
ARM64 maintainer Will Deacon suggested the pragmatic way out: Instead of fragile auto-detection logic baked into the kernel, Linux should simply expose dedicated boot parameters so the behavior can be driven explicitly from the outside.
That split the responsibility cleanly down the middle:
- In the Linux kernel: The existing
idle=nopparameter instructs the core CPU idle path to bypass WFI. Complementing that, upstream commit arm64: Add override for WFxT introducedarm64.nowfxtto explicitly disable kernel WFIT/WFxT handling. - In the m1n1 bootloader: Hardware detection lands where it belongs: in the bootloader. m1n1 inspects platform and CPU feature registers to identify affected bare-metal targets, appending
idle=nop arm64.nowfxtto the kernel command line automatically (see commit kboot: disable wfi/wfit if wfi loses state).
This doesn't fix the register-wiping behavior on the M4, but it fences it off cleanly:
Linux no longer crashes on boot, virtual machines remain completely unaffected, and the code merged into mainline Linux. For deeper power states, the Asahi project relies for now on its dedicated cpuidle-apple driver, while cleaner EFI/PSCI interfaces are developed for the future.
This is where kitchen-table hacking separates from real systems engineering: The real achievement isn't slapping a rough hack into assembly code. It's the discipline required to turn that hack into a maintainable solution that respects Linux's existing architecture.
What Apple Doesn't Document, the Community Re-Engineers
Apple's silicon architecture is technically impressive, efficient, and thoroughly optimized for its own software stack. But the M4 bring-up once again reveals what that walled garden means in the trenches: Linux support doesn't happen because someone opens a hardware datasheet and writes a clean device driver. It happens because engineers spend days poking blindly at registers until a serial line coughs up a single ASCII character.
From Apple's perspective, this makes complete sense.
Apple sells tightly integrated systems, not open developer boards. And it shows: As long as macOS and the XNU kernel run smoothly, Linux compatibility is nowhere on Apple's product roadmap. It is enough that Apple's own software understands the idiosyncrasies of Apple's own silicon.
For open operating systems, that makes life difficult.
Complaining about it doesn't move the needle. What projects like Asahi Linux and engineers like Yureka Lilian deliver isn't glamorous heroics; it is gritty, invisible grunt work: reverse engineering at the pain threshold, authoring custom bootloader stages, and painstakingly probing chips for which no public specification is available.
The real accomplishment isn't strong-arming Linux into booting on a stage demo with a couple of brittle hacks. The true craft lies in capturing these quirks cleanly enough to land them as upstream patches in mainline Linux—without breaking other ARM64 platforms or bogging down virtualized guests.
Walled-off platforms save hardware vendors significant support and documentation overhead. But that bill doesn't disappear; it simply gets forwarded downstream. And it's paid in full by anyone who insists on truly owning and controlling their hardware—choosing for themselves which operating system runs on it.
Why Sysadmins Should Care
In day-to-day operations, modern Linux usually feels domesticated. We boot pre-baked cloud images, let package managers resolve dependencies, orchestrate containers, roll out environments via Ansible, and click around polished desktop environments. The abstractions have grown so smooth that it is easy to forget how much raw mechanical friction hums underneath.
You get accustomed to the abstraction. We treat compute like electricity from a wall socket:
Power on, kernel loads, userspace runs.
The M4 bring-up is a blunt reminder that this comfort is an illusion.
Underneath every systemd unit, every container layer, and every polished desktop interface still lie the unforgiving low-level foundations: bootloaders mapping address windows, MMU page tables, interrupt controllers, CPU registers, and hardware timers.
When that plumbing breaks, high-level tooling won't save you. No modern installer rescues an OS when the processor zeros out its own registers after a sleep cycle. No container runtime spins up when an emergency debug routine dereferences unmapped virtual space. In that failure mode, there are no logs, no journalctl, and no emergency rescue shell. The machine is dead in the water.
For sysadmins, platform engineers, and anyone running infrastructure at scale, this is a healthy dose of perspective. You don't need to write ARM64 assembly or reverse-engineer proprietary Apple registers every day to be good at your job. But you should never lose respect for the bare-metal layers that make your daily routine possible.
Sound systems engineering often starts where automated abstractions tap out: when a hypervisor behaves erratically under bursty load, when a kernel hangs consistently after an update on specific silicon, or when kernel panics smell like memory controller bugs:
That is when foundational understanding of how kernel and silicon intersect actually matters.
If you see infrastructure only as a pile of configuration files, you're helpless the moment an abstraction layer fractures. But if you appreciate how deep the roots go—and the sheer amount of stubborn troubleshooting it takes to coax a machine into displaying a single ASCII character on a terminal—you look at your stack differently: with a bit more respect for the machine, and far less patience for marketing fairy tales about systems that supposedly "just work."
A Single Letter as a Milestone
At the end of the day, there is no turnkey desktop, no slick installer, and no graphics acceleration maxing out benchmark graphs. Booting Linux on an Apple M4 today lands you in a spartan shell. There is still real work left before the remaining platform peripherals reach the standard established on older M-series silicon with Asahi Linux.
Yet the initial, stubborn wall has crumbled.
From a single ASCII character—that lonely a shoved down a raw serial link by hand-rolled assembly—emerged a booting operating system. That first heartbeat opened the door to systematic bisection, which yielded targeted workarounds, and ultimately landed clean code inside mainline Linux. Not by magic, and certainly not through vendor benevolence, but by methodically dissecting problems nobody had ever documented.
Before Linux can feel comfortable on new silicon, it first has to prove that it is alive. On the Apple M4, that proof was a single a. Sometimes that is enough to get things moving.