Cache coherence in multi-core CPUs: how it is maintained and who controls it

Last update: March 6th 2026
  • Cache coherence ensures that all copies of the same data in different caches and in RAM remain consistent on multi-core systems.
  • The cache hierarchy with a shared last level simplifies consistency control and reduces direct accesses to main memory.
  • Coherence protocols use copy invalidation or update strategies, supported by states and control bits per cache line.
  • The compiler and operating system can supplement hardware consistency by inserting instructions and configuring memory for critical periods.

CPU cache coherence scheme

When you look at a diagram of any modern multi-core processor, the same pattern always appears: multiple cores, each with its own nearby caches, and a shared last-level cache that acts as a common point before reaching RAM. This arrangement is not accidental or a whim of the designers, but a direct response to a critical problem in parallel systems: cache coherence.

Without a robust consistency mechanism, each core could end up working with a different and outdated version of the same data in memory , which in a real-world program translates to subtle errors, unpredictable failures, and even system crashes. Therefore, understanding how this consistency is maintained—at both the hardware and software levels—is key to understanding the performance and stability of modern multi-core CPUs.

What is cache coherence: the terminal metaphor

real-time in electronic systems
Related articles:
Real-time electronic systems: fundamentals, planning, and applications

Imagine several people sitting in front of different terminals, all editing the same document stored on a central server . Each screen displays a copy of the file, and any changes one person makes are expected to be reflected immediately on everyone else's screens.

For this to work, there needs to be a synchronization mechanism that propagates document changes to all terminals, so that everyone always sees the same version. As long as this system is working, everything is fine: whoever modifies the text knows that everyone else will see the new version almost instantly.

Now imagine that the synchronization system suddenly fails. Each person continues editing, convinced they're working on the shared document, but in reality, each terminal is left with its own disconnected local copy . From that moment on, changes made by one person don't reach the others, and the document begins to diverge uncontrollably.

In the realm of computing, this is precisely what would happen if the CPU lacked a reliable consistency protocol: one core modifies data in memory, but the other cores continue reading an older version from their private caches . This creates fertile ground for serious logical errors, corrupted data, and undebugging behavior.

Cache coherence is therefore the set of mechanisms that ensure that, in a multi-core system, all copies of the same data distributed across the different caches and RAM maintain a consistent state . Even if multiple copies exist, the system must behave "as if" there were only one.

Cache hierarchy in multi-core CPUs

Caches and memory hierarchy in a multi-core CPU

CPU caches are small, very fast memories that hold copies of frequently used chunks of RAM . When the processor executes code, instead of continuously accessing (comparatively slow) RAM, it attempts to read from and write to the cache, drastically reducing latency.

The trick, of course, is that caches don't store the "official version" of the data, but only a temporary replica . Following the terminal metaphor, RAM would be the document on the server, while caches would be the local screens that display copies of certain parts of the file.

In a multi-core CPU, the design becomes more complex because each core typically has its own private Level 1 (L1) and even Level 2 (L2) caches . Above these, a shared Level 3 cache (for example) is added, located between the cores and the memory controller that provides access to RAM.

This shared cache is introduced because allowing all cores to directly and intensively access RAM would cause access conflicts, contention on the memory bus, and a significant performance drop . The last-level cache acts as a common "buffer" that reduces RAM accesses and centralizes much of the data traffic.

  PC with artificial intelligence: real differences compared to a traditional PC

Furthermore, many architectures organize caches inclusively: lines stored in levels close to the processor are also present in higher levels of the hierarchy . That is, a line that appears in L1 is also in L2 and, in turn, in L3. This has a very useful consequence for consistency: simply updating the lowest-level cache correctly is enough to control the state of the other levels without having to constantly access RAM.

Why last-level shared caching is key to consistency

Without this global last-level cache, each core would have to check for consistency directly against main memory . Every time a memory line in a private cache was modified, it would be necessary to check if other cores maintain a copy of that same line and, if so, update or invalidate it everywhere.

In a system with many cores, this workload of checks would result in a huge number of transactions to RAM , negating much of the benefit of having fast caches. By placing a shared cache between the cores and memory, the CPU can concentrate coherence control in a single intermediate location.

In many implementations, the caches at higher levels (further from the processor) contain copies of the lines present in the levels closer to the core . With this organization, the coherence protocol only needs to ensure that the last level is synchronized with main memory, and that the private levels of each core are synchronized with the level immediately above it.

This can be visualized as a kind of Russian nesting doll: the third-level cache includes the content of the second and first levels , the second level includes its own content and that of the first level, and the first level only knows its own lines. Thus, by controlling the "big doll" (the last level), the system can coordinate the rest more efficiently.

The result is that maintaining consistency becomes more economical in terms of design and memory traffic . Instead of forcing each core to constantly deal with RAM, the protocol operates on the shared cache and from there manages which lines should be updated or invalidated in the private caches.

Update methods: invalidation and updating of copies

A critical issue arises when two or more cores want to access, almost simultaneously, the same line of data that is replicated across multiple caches . In this context, consistency systems typically employ two fundamental strategies when handling writes.

The first method is based on invalidation. When a kernel needs to write to a specific cache line, the protocol invalidates any copies of that same line that may exist in the other caches . Only the kernel that is going to write keeps the line in a read- and write-enabled state; the others, if they want to use that data again, will have to reload the line from the higher level (or from memory) with the updated version.

The second strategy involves updating. In this case, when a kernel modifies a line, the system attempts to automatically propagate the new content to existing copies in the other caches . This way, all caches that stored that line receive the updated version without needing to invalidate and reload it later.

Each approach has its pros and cons. Invalidation is usually more efficient when writes are frequent because it avoids saturating the memory system with updates that other cores may not immediately need. Conversely, updating can be advantageous when many cores frequently read the same data that is modified relatively infrequently , as it reduces latency by not having to reload the line after each invalidation.

In either case, both methods utilize additional states and control bits in the cache lines. Each line typically includes information about whether its contents match those in RAM , and whether it is shared, modified, exclusive, reserved, etc., depending on the specific protocol (MESI, MOESI, MSI, etc.). This allows the hardware to make quick decisions about what to do when a read or write operation occurs on an already replicated line.

  Xiaomi 17 Max: Extreme Power and Unprecedented Autonomy

Checking consistency between caches and memory

Directly verifying the consistency between all cache levels of a CPU or GPU and main memory would be a mammoth task, both in terms of design complexity and performance cost. Therefore, modern systems organize this verification hierarchically.

The caches closest to the processor (L1, L2) are not usually connected directly to RAM, but to the next level of cache. This means that consistency is not validated against main memory at each level, but rather against the immediately higher level . This reduces the number of RAM accesses and simplifies the logic required at lower levels.

Ultimately, the comparison between cache contents and RAM contents is performed between the last-level cache and main memory . If this last level maintains a correct and consistent state, and each lower level maintains its consistency with the one above it, the entire hierarchy remains consistent without having to check each line against RAM repeatedly.

When a kernel writes to a cache line and changes its data, the state of that line is marked to indicate that it no longer exactly matches the copy stored in memory . From there, the protocol coordinates the update: it marks the corresponding copies in other caches as reserved or invalid and, when appropriate, writes the new content to the associated main memory line.

This cascading organization allows changes to propagate progressively from the kernel, which updates the data, to main memory, passing through each cache level in a controlled manner. In this way, maintaining consistency does not become an insurmountable bottleneck for the processor.

Hardware coherence versus software coherence

So far we have discussed consistency mechanisms that are mainly implemented in hardware: protocols, status bits, shared caches, etc. However, there is another approach that seeks to shift some of that complexity to software , specifically to the compiler and the operating system.

Software-based consistency schemes attempt to reduce the need for additional on-chip logic by analyzing code and making compile-time decisions . The idea is that if the compiler can deduce when and how certain shared data is accessed, it could, in many cases, prevent that data from being cached or explicitly manage its visibility.

This approach has a clear advantage: part of the workload shifts from runtime to compile-time resolution . Instead of the hardware detecting and handling all conflicts on the fly, the compiler attempts to anticipate them and generate code that avoids dangerous situations.

The downside is that static code analysis is limited, and therefore compilers tend to be conservative . This means that, to avoid violating consistency, they often make decisions that reduce the effectiveness of caches. If they suspect that some data might be problematic, they frequently prevent it from being cached or force synchronizations more frequently than strictly necessary.

Therefore, although these software schemes are attractive in theory, especially for simplifying hardware design, in practice they do not replace the coherence support integrated into the CPU itself , but rather complement it in some specific scenarios.

The role of the compiler in cache consistency

A key element of software-based consistency approaches is the role of the compiler. The compiler can perform a deep analysis of the code and determine which shared data structures might be unsafe for caching . Based on this, it marks these elements in a special way or adapts the code generation.

The simplest, and also the most conservative, approach is to prevent shared data variables from being cached . That is, each access to these variables forces an access to main memory or a non-cacheable area. This guarantees consistency, but misses many performance opportunities, because a shared structure can, in fact, be used privately during certain periods, or read-only at others.

  Supercomputing, AI and digital twins: a complete guide in Spanish

In reality, the consistency problem only arises during intervals when at least one process can write to the variable and another process can read it . Outside of these critical periods, the variable can be treated as being for the exclusive use of a single thread or even as an effective constant for a while, allowing it to be cached without issues.

The most advanced compilation strategies attempt to identify those "safe" periods during which the shared variable can be considered non-conflicting . To do this, the compiler analyzes execution paths, potential concurrent accesses, and synchronization patterns (locks, critical sections, etc.). Based on this analysis, it divides the variable's lifetime into phases: some suitable for caching, others requiring special handling.

During critical periods, when concurrent access with writes is detected, the compiler inserts additional instructions into the generated code to enforce cache consistency . These instructions may force cache flushes, memory reloads, memory barriers, or access to regions marked as non-cacheable, depending on the programming model and underlying architecture.

Relationship between compiler, operating system, and hardware

The phrase "the compiler inserts instructions into the generated code to enforce cache consistency" might lead one to think that the operating system reads these instructions as if they were high-level hints and, based on that, decides how to execute the program. In reality, the mechanism is somewhat different.

When the compiler adds these types of instructions, what it introduces into the binary are specific operations supported by the architecture or the runtime environment . For example, it can insert cache flushing instructions, memory barriers, special instructions to mark regions as non-cacheable, or calls to operating system services that configure memory attributes.

The operating system doesn't interpret these instructions as high-level "comments" or "hints" written by the compiler; it simply executes the machine code like any other . However, some of these instructions are designed to interact with the memory subsystem and cache management, thus changing how the CPU accesses certain data.

In other words, the compiler performs preliminary analysis and generates code that, when executed, produces the desired cache behavior . The operating system collaborates by establishing memory attributes (cacheable or non-cacheable areas, write policies, etc.) and providing synchronization primitives, but it is not "reading" special instructions in the sense of interpreting them semantically as a compiler would.

It can also happen that the hardware, upon seeing certain instructions, activates specific coherence or synchronization mechanisms . For example, fence or barrier instructions guarantee the order of memory access and enforce certain visibility effects across the cache hierarchy. In this case, there is a three-way collaboration: the compiler decides where to place these instructions, the operating system configures the execution environment, and the hardware implements the actual behavior at the cache and memory bus level.

Together, all these elements ensure that, even with multiple copies of the same data distributed across different caches and main memory, parallel programs run with a consistent memory model . Cache coherence, far from being a simple internal CPU detail, becomes a central component for multi-core systems to operate reliably and efficiently.

Understanding how cache hierarchy, hardware coherence protocols, and software support techniques combine makes it clearer why modern CPU designs share such a similar structure, and why a small failure in any of those mechanisms can trigger chaotic behavior in concurrent applications that depend entirely on all cores seeing the same data at the right time.