Skip to main content

sidecar/
main.rs

1// Copyright (c) Microsoft Corporation.
2// Licensed under the MIT License.
3
4#![cfg_attr(minimal_rt, no_std, no_main)]
5
6//! This crate implements the OpenHCL sidecar kernel. This is a kernel that runs
7//! along side the OpenHCL Linux kernel, operating on a subset of the virtual
8//! machine's CPUs.
9//!
10//! Sidecar is an x86_64 bare-metal payload packaged into an OpenHCL IGVM, not a
11//! host command-line program. The OpenHCL boot loader initializes it before
12//! Linux starts; non-`minimal_rt` builds contain only a stub that rejects
13//! execution. Developers normally build and consume it through `cargo xflowey
14//! build-igvm`.
15//!
16//! This is done to avoid needing to boot all CPUs into Linux, since this is
17//! very expensive for large VMs. Instead, most of the CPUs are run in the
18//! sidecar kernel, where they run a minimal dispatch loop. If a sidecar CPU
19//! hits a condition that it cannot handle locally (e.g., the guest OS attempts
20//! to access an emulated device), it will send a message to the main Linux
21//! kernel. One of the Linux CPUs can then handle the exit remotely, and/or
22//! convert the sidecar CPU to a Linux CPU.
23//!
24//! Similarly, if a Linux CPU needs to run code on a sidecar CPU (e.g., to run
25//! it as a target for device interrupts from the host), it can convert the
26//! sidecar CPU to a Linux CPU.
27//!
28//! Sidecar is modeled to Linux as a set of devices, one per node (a contiguous
29//! set of CPUs; this may or may not correspond to a NUMA node or CPU package).
30//! Each device has a single control page, used to communicate with the sidecar
31//! CPUs. Each CPU additionally has a command page, which is used to specify
32//! sidecar commands (e.g., run the VP, or get or set VP registers). These
33//! commands are in separate pages at least partially so that they can be
34//! operated on independently; the Linux kernel communicates with sidecar via
35//! control page, and the user-mode VMM communicates with the individual sidecar
36//! CPUs via the command pages.
37//!
38//! The sidecar kernel is a very simple kernel. It runs at a fixed virtual
39//! address (although it is still built with dynamic relocations). Each CPU has
40//! its own set of page tables (sharing some portion of them) so that they only
41//! map what they use. Each CPU is independent after boot; sidecar CPUs never
42//! communicate with each other and only communicate with Linux CPUs, via the
43//! Linux sidecar driver.
44//!
45//! The sidecar CPU runs a simple dispatch loop. It halts the processor, waiting
46//! for the control page to indicate that it should run (the sidecar driver
47//! sends an IPI when the control page is updated). It then reads a command from
48//! the command page and executes the command; if the command can run for an
49//! unbounded amount of time (e.g., the command to run the VP), then the driver
50//! can interrupt the command via another request on the control page (and
51//! another IPI).
52//!
53//! # Processor Startup
54//!
55//! The sidecar kernel is initialized by a single bootstrap processor (BSP),
56//! which is typically VP 0. This initialization happens during the boot shim
57//! phase, before the Linux kernel starts. The BSP (which will later become a
58//! Linux CPU) calls into the sidecar kernel to perform all global initialization
59//! tasks: copying the hypercall page, setting up the IDT, initializing control
60//! pages for each node, and preparing page tables and per-CPU state for all
61//! application processors (APs).
62//!
63//! After the BSP completes its initialization work, it begins starting APs. The
64//! startup process uses a fan-out pattern to minimize total boot time: the BSP
65//! starts the first few APs, and then each newly-started AP immediately helps
66//! start additional APs. This creates an exponential growth in the number of
67//! CPUs actively participating in the boot process.
68//!
69//! Concurrency during startup is managed through atomic operations on a
70//! per-node `next_vp` counter. Each CPU (whether BSP or AP) atomically
71//! increments this counter to claim the next VP index to start within that
72//! node. This ensures that each VP is started exactly once without requiring
73//! locks or complex coordination. The startup fan-out continues until all VPs
74//! in all nodes have been started (or skipped if marked as REMOVED).
75//!
76//! Note that the first VP in each NUMA node is typically reserved for the Linux
77//! kernel and does not run the sidecar kernel. The sidecar startup logic
78//! accounts for this by initializing the `next_vp` counter to 1 for each node,
79//! effectively skipping the base VP (index 0) of that node.
80//!
81//! Each CPU's page tables include a mapping for its node's control page at a
82//! fixed virtual address (PTE_CONTROL_PAGE). This is set up during AP
83//! initialization via the `init_ap` function, which builds the per-CPU page
84//! table hierarchy and maps the control page at the same virtual address for
85//! all CPUs in the node. This allows each CPU to access its control page
86//! without knowing its physical address, and ensures that all CPUs in a node
87//! see the same control page data (since they all map the same physical page).
88//! The control page mapping is read-write from the sidecar's perspective, as
89//! the sidecar needs to update status fields (like `cpu_status` and
90//! `needs_attention`) using atomic operations.
91//!
92//! As of this writing, sidecar only supports x86_64, without hardware
93//! isolation.
94
95mod arch;
96
97#[cfg(not(minimal_rt))]
98fn main() {
99    panic!("must build with MINIMAL_RT_BUILD=1")
100}