dm-zoned.rst 8.3 KB

123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180181182183184185186187188189190191192193194
  1. ========
  2. dm-zoned
  3. ========
  4. The dm-zoned device mapper target exposes a zoned block device (ZBC and
  5. ZAC compliant devices) as a regular block device without any write
  6. pattern constraints. In effect, it implements a drive-managed zoned
  7. block device which hides from the user (a file system or an application
  8. doing raw block device accesses) the sequential write constraints of
  9. host-managed zoned block devices and can mitigate the potential
  10. device-side performance degradation due to excessive random writes on
  11. host-aware zoned block devices.
  12. For a more detailed description of the zoned block device models and
  13. their constraints see (for SCSI devices):
  14. https://www.t10.org/drafts.htm#ZBC_Family
  15. and (for ATA devices):
  16. http://www.t13.org/Documents/UploadedDocuments/docs2015/di537r05-Zoned_Device_ATA_Command_Set_ZAC.pdf
  17. The dm-zoned implementation is simple and minimizes system overhead (CPU
  18. and memory usage as well as storage capacity loss). For a 10TB
  19. host-managed disk with 256 MB zones, dm-zoned memory usage per disk
  20. instance is at most 4.5 MB and as little as 5 zones will be used
  21. internally for storing metadata and performing reclaim operations.
  22. dm-zoned target devices are formatted and checked using the dmzadm
  23. utility available at:
  24. https://github.com/hgst/dm-zoned-tools
  25. Algorithm
  26. =========
  27. dm-zoned implements an on-disk buffering scheme to handle non-sequential
  28. write accesses to the sequential zones of a zoned block device.
  29. Conventional zones are used for caching as well as for storing internal
  30. metadata. It can also use a regular block device together with the zoned
  31. block device; in that case the regular block device will be split logically
  32. in zones with the same size as the zoned block device. These zones will be
  33. placed in front of the zones from the zoned block device and will be handled
  34. just like conventional zones.
  35. The zones of the device(s) are separated into 2 types:
  36. 1) Metadata zones: these are conventional zones used to store metadata.
  37. Metadata zones are not reported as usable capacity to the user.
  38. 2) Data zones: all remaining zones, the vast majority of which will be
  39. sequential zones used exclusively to store user data. The conventional
  40. zones of the device may be used also for buffering user random writes.
  41. Data in these zones may be directly mapped to the conventional zone, but
  42. later moved to a sequential zone so that the conventional zone can be
  43. reused for buffering incoming random writes.
  44. dm-zoned exposes a logical device with a sector size of 4096 bytes,
  45. irrespective of the physical sector size of the backend zoned block
  46. device being used. This allows reducing the amount of metadata needed to
  47. manage valid blocks (blocks written).
  48. The on-disk metadata format is as follows:
  49. 1) The first block of the first conventional zone found contains the
  50. super block which describes the on disk amount and position of metadata
  51. blocks.
  52. 2) Following the super block, a set of blocks is used to describe the
  53. mapping of the logical device blocks. The mapping is done per chunk of
  54. blocks, with the chunk size equal to the zoned block device size. The
  55. mapping table is indexed by chunk number and each mapping entry
  56. indicates the zone number of the device storing the chunk of data. Each
  57. mapping entry may also indicate if the zone number of a conventional
  58. zone used to buffer random modification to the data zone.
  59. 3) A set of blocks used to store bitmaps indicating the validity of
  60. blocks in the data zones follows the mapping table. A valid block is
  61. defined as a block that was written and not discarded. For a buffered
  62. data chunk, a block is always valid only in the data zone mapping the
  63. chunk or in the buffer zone of the chunk.
  64. For a logical chunk mapped to a conventional zone, all write operations
  65. are processed by directly writing to the zone. If the mapping zone is a
  66. sequential zone, the write operation is processed directly only if the
  67. write offset within the logical chunk is equal to the write pointer
  68. offset within of the sequential data zone (i.e. the write operation is
  69. aligned on the zone write pointer). Otherwise, write operations are
  70. processed indirectly using a buffer zone. In that case, an unused
  71. conventional zone is allocated and assigned to the chunk being
  72. accessed. Writing a block to the buffer zone of a chunk will
  73. automatically invalidate the same block in the sequential zone mapping
  74. the chunk. If all blocks of the sequential zone become invalid, the zone
  75. is freed and the chunk buffer zone becomes the primary zone mapping the
  76. chunk, resulting in native random write performance similar to a regular
  77. block device.
  78. Read operations are processed according to the block validity
  79. information provided by the bitmaps. Valid blocks are read either from
  80. the sequential zone mapping a chunk, or if the chunk is buffered, from
  81. the buffer zone assigned. If the accessed chunk has no mapping, or the
  82. accessed blocks are invalid, the read buffer is zeroed and the read
  83. operation terminated.
  84. After some time, the limited number of conventional zones available may
  85. be exhausted (all used to map chunks or buffer sequential zones) and
  86. unaligned writes to unbuffered chunks become impossible. To avoid this
  87. situation, a reclaim process regularly scans used conventional zones and
  88. tries to reclaim the least recently used zones by copying the valid
  89. blocks of the buffer zone to a free sequential zone. Once the copy
  90. completes, the chunk mapping is updated to point to the sequential zone
  91. and the buffer zone freed for reuse.
  92. Metadata Protection
  93. ===================
  94. To protect metadata against corruption in case of sudden power loss or
  95. system crash, 2 sets of metadata zones are used. One set, the primary
  96. set, is used as the main metadata region, while the secondary set is
  97. used as a staging area. Modified metadata is first written to the
  98. secondary set and validated by updating the super block in the secondary
  99. set, a generation counter is used to indicate that this set contains the
  100. newest metadata. Once this operation completes, in place of metadata
  101. block updates can be done in the primary metadata set. This ensures that
  102. one of the set is always consistent (all modifications committed or none
  103. at all). Flush operations are used as a commit point. Upon reception of
  104. a flush request, metadata modification activity is temporarily blocked
  105. (for both incoming BIO processing and reclaim process) and all dirty
  106. metadata blocks are staged and updated. Normal operation is then
  107. resumed. Flushing metadata thus only temporarily delays write and
  108. discard requests. Read requests can be processed concurrently while
  109. metadata flush is being executed.
  110. If a regular device is used in conjunction with the zoned block device,
  111. a third set of metadata (without the zone bitmaps) is written to the
  112. start of the zoned block device. This metadata has a generation counter of
  113. '0' and will never be updated during normal operation; it just serves for
  114. identification purposes. The first and second copy of the metadata
  115. are located at the start of the regular block device.
  116. Usage
  117. =====
  118. A zoned block device must first be formatted using the dmzadm tool. This
  119. will analyze the device zone configuration, determine where to place the
  120. metadata sets on the device and initialize the metadata sets.
  121. Ex::
  122. dmzadm --format /dev/sdxx
  123. If two drives are to be used, both devices must be specified, with the
  124. regular block device as the first device.
  125. Ex::
  126. dmzadm --format /dev/sdxx /dev/sdyy
  127. Formatted device(s) can be started with the dmzadm utility, too.:
  128. Ex::
  129. dmzadm --start /dev/sdxx /dev/sdyy
  130. Information about the internal layout and current usage of the zones can
  131. be obtained with the 'status' callback from dmsetup:
  132. Ex::
  133. dmsetup status /dev/dm-X
  134. will return a line
  135. 0 <size> zoned <nr_zones> zones <nr_unmap_rnd>/<nr_rnd> random <nr_unmap_seq>/<nr_seq> sequential
  136. where <nr_zones> is the total number of zones, <nr_unmap_rnd> is the number
  137. of unmapped (ie free) random zones, <nr_rnd> the total number of zones,
  138. <nr_unmap_seq> the number of unmapped sequential zones, and <nr_seq> the
  139. total number of sequential zones.
  140. Normally the reclaim process will be started once there are less than 50
  141. percent free random zones. In order to start the reclaim process manually
  142. even before reaching this threshold the 'dmsetup message' function can be
  143. used:
  144. Ex::
  145. dmsetup message /dev/dm-X 0 reclaim
  146. will start the reclaim process and random zones will be moved to sequential
  147. zones.