memory.rst 7.4 KB

123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165
  1. ==============================
  2. Memory Layout on AArch64 Linux
  3. ==============================
  4. Author: Catalin Marinas <catalin.marinas@arm.com>
  5. This document describes the virtual memory layout used by the AArch64
  6. Linux kernel. The architecture allows up to 4 levels of translation
  7. tables with a 4KB page size and up to 3 levels with a 64KB page size.
  8. AArch64 Linux uses either 3 levels or 4 levels of translation tables
  9. with the 4KB page configuration, allowing 39-bit (512GB) or 48-bit
  10. (256TB) virtual addresses, respectively, for both user and kernel. With
  11. 64KB pages, only 2 levels of translation tables, allowing 42-bit (4TB)
  12. virtual address, are used but the memory layout is the same.
  13. ARMv8.2 adds optional support for Large Virtual Address space. This is
  14. only available when running with a 64KB page size and expands the
  15. number of descriptors in the first level of translation.
  16. TTBRx selection is given by bit 55 of the virtual address. The
  17. swapper_pg_dir contains only kernel (global) mappings while the user pgd
  18. contains only user (non-global) mappings. The swapper_pg_dir address is
  19. written to TTBR1 and never written to TTBR0.
  20. AArch64 Linux memory layout with 4KB pages + 4 levels (48-bit)::
  21. Start End Size Use
  22. -----------------------------------------------------------------------
  23. 0000000000000000 0000ffffffffffff 256TB user
  24. ffff000000000000 ffff7fffffffffff 128TB kernel logical memory map
  25. [ffff600000000000 ffff7fffffffffff] 32TB [kasan shadow region]
  26. ffff800000000000 ffff80007fffffff 2GB modules
  27. ffff800080000000 fffffbffefffffff 124TB vmalloc
  28. fffffbfff0000000 fffffbfffdffffff 224MB fixed mappings (top down)
  29. fffffbfffe000000 fffffbfffe7fffff 8MB [guard region]
  30. fffffbfffe800000 fffffbffff7fffff 16MB PCI I/O space
  31. fffffbffff800000 fffffbffffffffff 8MB [guard region]
  32. fffffc0000000000 fffffdffffffffff 2TB vmemmap
  33. fffffe0000000000 ffffffffffffffff 2TB [guard region]
  34. AArch64 Linux memory layout with 64KB pages + 3 levels (52-bit with HW support)::
  35. Start End Size Use
  36. -----------------------------------------------------------------------
  37. 0000000000000000 000fffffffffffff 4PB user
  38. fff0000000000000 ffff7fffffffffff ~4PB kernel logical memory map
  39. [fffd800000000000 ffff7fffffffffff] 512TB [kasan shadow region]
  40. ffff800000000000 ffff80007fffffff 2GB modules
  41. ffff800080000000 fffffbffefffffff 124TB vmalloc
  42. fffffbfff0000000 fffffbfffdffffff 224MB fixed mappings (top down)
  43. fffffbfffe000000 fffffbfffe7fffff 8MB [guard region]
  44. fffffbfffe800000 fffffbffff7fffff 16MB PCI I/O space
  45. fffffbffff800000 fffffbffffffffff 8MB [guard region]
  46. fffffc0000000000 ffffffdfffffffff ~4TB vmemmap
  47. ffffffe000000000 ffffffffffffffff 128GB [guard region]
  48. Translation table lookup with 4KB pages::
  49. +--------+--------+--------+--------+--------+--------+--------+--------+
  50. |63 56|55 48|47 40|39 32|31 24|23 16|15 8|7 0|
  51. +--------+--------+--------+--------+--------+--------+--------+--------+
  52. | | | | | |
  53. | | | | | v
  54. | | | | | [11:0] in-page offset
  55. | | | | +-> [20:12] L3 index
  56. | | | +-----------> [29:21] L2 index
  57. | | +---------------------> [38:30] L1 index
  58. | +-------------------------------> [47:39] L0 index
  59. +----------------------------------------> [55] TTBR0/1
  60. Translation table lookup with 64KB pages::
  61. +--------+--------+--------+--------+--------+--------+--------+--------+
  62. |63 56|55 48|47 40|39 32|31 24|23 16|15 8|7 0|
  63. +--------+--------+--------+--------+--------+--------+--------+--------+
  64. | | | | |
  65. | | | | v
  66. | | | | [15:0] in-page offset
  67. | | | +----------> [28:16] L3 index
  68. | | +--------------------------> [41:29] L2 index
  69. | +-------------------------------> [47:42] L1 index (48-bit)
  70. | [51:42] L1 index (52-bit)
  71. +----------------------------------------> [55] TTBR0/1
  72. When using KVM without the Virtualization Host Extensions, the
  73. hypervisor maps kernel pages in EL2 at a fixed (and potentially
  74. random) offset from the linear mapping. See the kern_hyp_va macro and
  75. kvm_update_va_mask function for more details. MMIO devices such as
  76. GICv2 gets mapped next to the HYP idmap page, as do vectors when
  77. ARM64_SPECTRE_V3A is enabled for particular CPUs.
  78. When using KVM with the Virtualization Host Extensions, no additional
  79. mappings are created, since the host kernel runs directly in EL2.
  80. 52-bit VA support in the kernel
  81. -------------------------------
  82. If the ARMv8.2-LVA optional feature is present, and we are running
  83. with a 64KB page size; then it is possible to use 52-bits of address
  84. space for both userspace and kernel addresses. However, any kernel
  85. binary that supports 52-bit must also be able to fall back to 48-bit
  86. at early boot time if the hardware feature is not present.
  87. This fallback mechanism necessitates the kernel .text to be in the
  88. higher addresses such that they are invariant to 48/52-bit VAs. Due
  89. to the kasan shadow being a fraction of the entire kernel VA space,
  90. the end of the kasan shadow must also be in the higher half of the
  91. kernel VA space for both 48/52-bit. (Switching from 48-bit to 52-bit,
  92. the end of the kasan shadow is invariant and dependent on ~0UL,
  93. whilst the start address will "grow" towards the lower addresses).
  94. In order to optimise phys_to_virt and virt_to_phys, the PAGE_OFFSET
  95. is kept constant at 0xFFF0000000000000 (corresponding to 52-bit),
  96. this obviates the need for an extra variable read. The physvirt
  97. offset and vmemmap offsets are computed at early boot to enable
  98. this logic.
  99. As a single binary will need to support both 48-bit and 52-bit VA
  100. spaces, the VMEMMAP must be sized large enough for 52-bit VAs and
  101. also must be sized large enough to accommodate a fixed PAGE_OFFSET.
  102. Most code in the kernel should not need to consider the VA_BITS, for
  103. code that does need to know the VA size the variables are
  104. defined as follows:
  105. VA_BITS constant the *maximum* VA space size
  106. VA_BITS_MIN constant the *minimum* VA space size
  107. vabits_actual variable the *actual* VA space size
  108. Maximum and minimum sizes can be useful to ensure that buffers are
  109. sized large enough or that addresses are positioned close enough for
  110. the "worst" case.
  111. 52-bit userspace VAs
  112. --------------------
  113. To maintain compatibility with software that relies on the ARMv8.0
  114. VA space maximum size of 48-bits, the kernel will, by default,
  115. return virtual addresses to userspace from a 48-bit range.
  116. Software can "opt-in" to receiving VAs from a 52-bit space by
  117. specifying an mmap hint parameter that is larger than 48-bit.
  118. For example:
  119. .. code-block:: c
  120. maybe_high_address = mmap(~0UL, size, prot, flags,...);
  121. It is also possible to build a debug kernel that returns addresses
  122. from a 52-bit space by enabling the following kernel config options:
  123. .. code-block:: sh
  124. CONFIG_EXPERT=y && CONFIG_ARM64_FORCE_52BIT=y
  125. Note that this option is only intended for debugging applications
  126. and should not be used in production.