]> git-server-git.apps.pok.os.sepia.ceph.com Git - ceph.git/log
ceph.git
3 weeks agoscript/buildcontainer-setup: fix package install on debian trixie 70128/head
David Galloway [Fri, 10 Jul 2026 20:12:01 +0000 (16:12 -0400)]
script/buildcontainer-setup: fix package install on debian trixie

Debian removed software-properties-common from the archive in trixie,
so the flat apt-get install list fails there. The package was only
needed to provide add-apt-repository for llvm.sh, which installs
clang-19 from apt.llvm.org. Trixie ships clang-19 natively, so install
it from the distro instead; run-make.sh's prepare() then finds clang-19
and skips llvm.sh entirely. Ubuntu and older Debian releases still have
software-properties-common and keep the previous behavior.

Fixes: https://tracker.ceph.com/issues/78111
Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit b5697987ae86f51d7f16a1409ce9f2b542f3de52)

4 weeks agoMerge pull request #69586 from mheler/wip-77494-umbrella
Neha Ojha [Wed, 1 Jul 2026 17:25:51 +0000 (10:25 -0700)]
Merge pull request #69586 from mheler/wip-77494-umbrella

umbrella: mon: implement mon backup mechanism

Reviewed-by: Kefu Chai <tchaikov@gmail.com>
Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
4 weeks agoMerge pull request #69549 from sseshasa/wip-77445-umbrella
Neha Ojha [Wed, 1 Jul 2026 17:24:06 +0000 (10:24 -0700)]
Merge pull request #69549 from sseshasa/wip-77445-umbrella

umbrella: mgr/DaemonServer: Aggregate and globally sort OSDs for ok-to-upgrade

Reviewed-by: Nitzan Mordechai <nmordech@redhat.com>
4 weeks agoMerge pull request #69707 from guits/wip-77645-umbrella
Neha Ojha [Tue, 30 Jun 2026 20:45:44 +0000 (13:45 -0700)]
Merge pull request #69707 from guits/wip-77645-umbrella

umbrella: node-proxy: atollon hardware monitoring (FCM stats, temperatures, fan speed..)

Reviewed-by: Afreen Misbah <afreen@ibm.com>
Reviewed-by: Adam King <adking@redhat.com>
4 weeks agoMerge pull request #69718 from adamemerson/wip-76995-umbrella
Neha Ojha [Tue, 30 Jun 2026 15:22:32 +0000 (08:22 -0700)]
Merge pull request #69718 from adamemerson/wip-76995-umbrella

umbrella: Reapply "qa/rgw/crypt: disable failing kmip testing"

Reviewed-by: Casey Bodley <cbodley@redhat.com>
4 weeks agoMerge pull request #69606 from mheler/wip-77528-umbrella
Neha Ojha [Tue, 30 Jun 2026 14:53:30 +0000 (07:53 -0700)]
Merge pull request #69606 from mheler/wip-77528-umbrella

umbrella: rgw/restore: run shard hash through HASH_PRIME

Reviewed-by: Adam Emerson <aemerson@redhat.com>
Reviewed-by: Soumya Koduri <skoduri@redhat.com>
4 weeks agoMerge pull request #69683 from mheler/wip-77614-umbrella
Neha Ojha [Tue, 30 Jun 2026 00:42:52 +0000 (17:42 -0700)]
Merge pull request #69683 from mheler/wip-77614-umbrella

umbrella: rgw: fix AES-256-GCM key/IV reuse on multipart part re-upload

Reviewed-by: Adam Emerson <aemerson@redhat.com>
4 weeks agoMerge pull request #69716 from adamemerson/wip-umbrella-include-s3tests
Neha Ojha [Tue, 30 Jun 2026 00:41:54 +0000 (17:41 -0700)]
Merge pull request #69716 from adamemerson/wip-umbrella-include-s3tests

umbrella: Include s3-tests in Ceph repo

Reviewed-by: Casey Bodley <cbodley@redhat.com>
4 weeks agoMerge pull request #69678 from rhcs-dashboard/wip-77518-umbrella
Neha Ojha [Mon, 29 Jun 2026 21:06:22 +0000 (14:06 -0700)]
Merge pull request #69678 from rhcs-dashboard/wip-77518-umbrella

umbrella: mgr/dashboard : Support wildcard sans and zonegroup hostnames

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 weeks agoMerge pull request #69698 from rhcs-dashboard/wip-77507-umbrella
Neha Ojha [Mon, 29 Jun 2026 21:05:53 +0000 (14:05 -0700)]
Merge pull request #69698 from rhcs-dashboard/wip-77507-umbrella

umbrella: mgr/dashboard: align RGW role management with Carbon and fix API routing

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 weeks agoMerge pull request #69704 from rhcs-dashboard/wip-77371-umbrella
Neha Ojha [Mon, 29 Jun 2026 21:05:22 +0000 (14:05 -0700)]
Merge pull request #69704 from rhcs-dashboard/wip-77371-umbrella

umbrella: mgr/dashboard: fix zone creation in rgw service creation form

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 weeks agoMerge pull request #69714 from afreen23/wip-77662-umbrella
Neha Ojha [Mon, 29 Jun 2026 21:00:57 +0000 (14:00 -0700)]
Merge pull request #69714 from afreen23/wip-77662-umbrella

umbrella: mgr/dashboard: fix bind address regression from CherryPy isolation

Reviewed-by: Nizamudeen A <nia@redhat.com>
Reviewed-by: Puja Shahu <pshahu@redhat.com>
4 weeks agoMerge pull request #69700 from guits/wip-77637-umbrella
Neha Ojha [Mon, 29 Jun 2026 20:53:13 +0000 (13:53 -0700)]
Merge pull request #69700 from guits/wip-77637-umbrella

umbrella: ceph-volume: skip internal raid mirror LVs in inventory

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 weeks agoMerge pull request #69719 from aainscow/wip-77668-umbrella
Neha Ojha [Mon, 29 Jun 2026 20:42:44 +0000 (13:42 -0700)]
Merge pull request #69719 from aainscow/wip-77668-umbrella

umbrella: osd/ECTransaction: fix truncate+write planning for EC shard sizes

Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
4 weeks agoceph-volume: skip internal raid mirror LVs in inventory 69700/head
Guillaume Abrioux [Thu, 18 Jun 2026 05:19:43 +0000 (07:19 +0200)]
ceph-volume: skip internal raid mirror LVs in inventory

ceph-volume inventory started including all LVM mapper devices after
c06bee965f1. On hosts with raid mirrored system volumes, that pulls in
hidden legs like var_rmeta_0 which have no /dev/vg/lv node and makes
cephadm's ceph-volume inventory call fail.

Skip those internal LVs in get_devices() and avoid rewriting the device
path to a missing lv_path in Device._parse().

Fixces: https://tracker.ceph.com/issues/77486

Signed-off-by: Guillaume Abrioux <gabrioux@ibm.com>
(cherry picked from commit daec7e125a167bb8fd0fb9faaa04879f0bb215d2)

5 weeks agonode-proxy: expose FCM stats 69707/head
Guillaume Abrioux [Thu, 18 Jun 2026 14:54:38 +0000 (16:54 +0200)]
node-proxy: expose FCM stats

Collect FCM stats locally from NVMe drives (vendor log page 0xCA)
and expose them via node-proxy, the cephadm agent, and
`ceph orch hardware status --category fcm`.

Introduce a node backend that aggregates Redfish data with
node-local collectors, since FCM metrics are not available
from the BMC.

Fixes: https://tracker.ceph.com/issues/77521
Signed-off-by: Guillaume Abrioux <gabrioux@ibm.com>
(cherry picked from commit e982412a7a1f1ec6fc2fe70d11e839365b50aa91)

5 weeks agonode-proxy: rename firmwares to firmware with legacy aliases
Guillaume Abrioux [Mon, 15 Jun 2026 13:18:48 +0000 (15:18 +0200)]
node-proxy: rename firmwares to firmware with legacy aliases

Let's use 'firmware' as the standard name in node-proxy, cephadm,
and orch hardware status.

We still accept 'firmwares' as a deprecated alias and read legacy
cache payloads transparently for backward compatibility.

Fixes: https://tracker.ceph.com/issues/77410
Signed-off-by: Guillaume Abrioux <gabrioux@ibm.com>
(cherry picked from commit 1cd4b32d72a4a14cbd797187e791e038422f4fc3)

5 weeks agonode-proxy: override with atollon specific
Guillaume Abrioux [Wed, 10 Jun 2026 10:51:47 +0000 (12:51 +0200)]
node-proxy: override with atollon specific

This adds AtollonSystem for AMI/Atollon BMCs:
- memory Id mapping,
- storage description fixes,
- StorageControllers based drive enrichment

It also wires vendor selection through cephadm (hw_monitoring_vendor)
and show slot/firmware in hardware status storage output.

Fixes: https://tracker.ceph.com/issues/77408
Signed-off-by: Guillaume Abrioux <gabrioux@ibm.com>
(cherry picked from commit 9b3bfd171afe6714cb23dc0fe55bacad671a94dc)

5 weeks agonode-proxy: add temperatures and fan speed
Guillaume Abrioux [Wed, 10 Jun 2026 10:51:14 +0000 (12:51 +0200)]
node-proxy: add temperatures and fan speed

This adds the temperatures category and fan speed information.

Fixes: https://tracker.ceph.com/issues/77408
Signed-off-by: Guillaume Abrioux <gabrioux@ibm.com>
(cherry picked from commit deabf21144054ada7e6eb64d35c38adc44db130a)

5 weeks agomgr/dashboard : Support wildcard sans and zonegroup hostnames 69678/head
Abhishek Desai [Tue, 26 May 2026 07:48:40 +0000 (13:18 +0530)]
mgr/dashboard : Support wildcard sans and zonegroup hostnames
fixes : https://tracker.ceph.com/issues/76795
Signed-off-by: Abhishek Desai <abhishek.desai1@ibm.com>
(cherry picked from commit eee6a15dcd845914595dbb1d76470bd4b73947c1)

5 weeks agomgr/dashboard: align RGW role management with Carbon and fix API routing 69698/head
Sagar Gopale [Wed, 10 Jun 2026 13:29:50 +0000 (18:59 +0530)]
mgr/dashboard: align RGW role management with Carbon and fix API routing

Fixes: https://tracker.ceph.com/issues/77328
Signed-off-by: Sagar Gopale <sagar.gopale@ibm.com>
(cherry picked from commit 1d3e33140b8f694e75a14f6530ec54712b984179)

5 weeks agomgr/dashboard: Remove global RGW tenant Roles tab and decommission routes
Sagar Gopale [Tue, 9 Jun 2026 08:03:07 +0000 (13:33 +0530)]
mgr/dashboard: Remove global RGW tenant Roles tab and decommission routes

Fixes: https://tracker.ceph.com/issues/77262
Signed-off-by: Sagar Gopale <sagar.gopale@ibm.com>
(cherry picked from commit 6dadd2f0309c0c016225a6891e814388d75aecd5)

5 weeks agomgr/dashboard: fix zone creation in rgw service creation form 69704/head
Aashish Sharma [Thu, 11 Jun 2026 09:19:04 +0000 (14:49 +0530)]
mgr/dashboard: fix zone creation in rgw service creation form

The zone creation request from the rgw service creation form was missing
the tier_type, sync_from and sync_from_all properties as a result the
zone creation was failing. This PR tends to fix this issue.

Fixes: https://tracker.ceph.com/issues/77263
Signed-off-by: Aashish Sharma <aasharma@redhat.com>
(cherry picked from commit 95a8c70be86569ef19f3d5c1b11f7cebed682db5)

5 weeks agomgr/dashboard: fix bind address regression from CherryPy isolation wip-77662-umbrella 69714/head
Afreen Misbah [Tue, 23 Jun 2026 13:30:22 +0000 (19:00 +0530)]
mgr/dashboard: fix bind address regression from CherryPy isolation

The CherryPy isolation refactor (PR #67227) accidentally changed the
dashboard bind address from wildcard (*:8443) to mon_ip:8443. The
get_mgr_ip() replacement was originally only for URI generation, but
the refactor passed the mutated address to CherryPyMgr.mount() as the
actual socket bind address.

This breaks the management gateway when its VIP is not on the same
interface as mon_ip, as the dashboard becomes unreachable on other
interfaces.

Preserve the original wildcard address for binding and only use
get_mgr_ip() for the advertised URI. Add regression test to prevent
future confusion between bind_addr and server_addr.

Fixes: https://tracker.ceph.com/issues/77491
Signed-off-by: Afreen Misbah <afreen@ibm.com>
(cherry picked from commit f45a1bfc93596f17370ee38ce241ae8123d81e8e)

5 weeks agoMerge pull request #69710 from tchaikov/wip-77609-umbrella
Neha Ojha [Fri, 26 Jun 2026 15:13:20 +0000 (08:13 -0700)]
Merge pull request #69710 from tchaikov/wip-77609-umbrella

umbrella: mgr/dashboard: skip the table when an nvmeof cli result has no columns

Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
Reviewed-by: Anthony D'Atri <anthony.datri@gmail.com>
Reviewed-by: Afreen Misbah <afreen23.git@gmail.com>
Reviewed-by: Devika Babrekar <devika.babrekar@ibm.com>
5 weeks agopython-common/cryptotools: stop using the removed X509Req API 69710/head
Kefu Chai [Sat, 13 Jun 2026 01:50:09 +0000 (09:50 +0800)]
python-common/cryptotools: stop using the removed X509Req API

pyOpenSSL deprecated OpenSSL.crypto.X509Req in 24.2.0 (2024-07-20) and
removed it in 26.3.0 (2026-06-12). as we don't pin pyopenssl, CI picked
up the new release, and create_self_signed_cert() started failing with:

  AttributeError: module 'OpenSSL.crypto' has no attribute 'X509Req'

this took down run-tox-mgr, run-tox-mgr-dashboard-py3 and the mypy check.

we only used X509Req to build a subject name and then copied it into the
X509 cert. so drop it, and set the subject on the cert directly. the
resulting cert stays the same: subject from dname, issuer set to the same
subject, self-signed.

Fixes: https://tracker.ceph.com/issues/77391
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
(cherry picked from commit 1dda56b1a00a6cbf520d932332ff097716ab256e)

5 weeks agoosd/ECTransaction: fix truncate+write planning for EC shard sizes 69719/head
Alex Ainscow [Tue, 9 Jun 2026 14:45:25 +0000 (15:45 +0100)]
osd/ECTransaction: fix truncate+write planning for EC shard sizes

There are multiple problems fixed here, which are caused by operations
which perform a truncate-then-write in a single transaction.

NOTE: The only known scenario for these operations are CLS-sparsify operations.
I recommend backporting these changes to Tentacle, however, as they are
regressions in the RADOS API which can lead to data corruption.

PROBLEM: Cache not invalidated correctly:
If projected_size >= orig_size, the invalidate_cache flag is not set in the plan
This means there is potentially data in the RMW extent cache, which
may cause data corruption in subsequent writes (although not the write
being made). No actual use case for this operation is known, so it may not be as serious as
it sounds.

FIX: Invalidate cache on any truncate or "delete_first"

PROBLEM: Parity writes permitted with truncates:
Currently parity delta writes do not support operations which may have
truncates. However, if the object ended up the same size as pre-truncate, a
parity delta write may have been attemtped. This may lead to data corruption.

FIX: Block parity writes in this case.
NOTE: It is not worth the development effort to support PDW in this scenario, as
performance benefit would be minimal overall.

PROBLEM: RMW Reads can occur beyond first truncate point.
Such reads reflect the pre-truncate data (which is incorrect) and be
preserved, even if the invalidate cache flag is set (since the cache is
invalidated before the reads). This can lead to similar corruption as
the earlier "cache not invalidated"

FIX: trim the read to the lower truncate size.

PROBLEM:
When performing a truncate, the parity shards may need to be updated to
reflect truncates on other shards. These truncate writes in the plan were
overriding, rather than adding to, other writes in the op.  The result was
a potentially truncated coding shard.  This leads to assertions on reads.

FIX: Replace = with insert()

PROBLEM: Some shards not set to correct size on truncate-then-write
If projected_size is smaller or equal to the original size, then the
code which attempts to correctly size a shard will not run.  However if
an operation performs a truncate then a partial write and does not
write to a particular shard AND the shard ends up smaller, then this
can lead to an incorrectly sized shard. This leads to assertion on reads.

FIX: Execute the shard-resize code on all truncates.

AI assistance was mainly used to write unit test. However, I cannot rule out a
contribution to the simple fixes found in this commit, so out of caution,
I place the Assisted-by tag.

Fixes: https://tracker.ceph.com/issues/77276
Signed-off-by: Alex Ainscow <aainscow@uk.ibm.com>
Assisted-by: IBM-Bob:ClaudeSonnet/GPT
(cherry picked from commit 51d8c5c489ba3e664209fb3316f8d6e03e257e28)

5 weeks agoReapply "qa/rgw/crypt: disable failing kmip testing" 69718/head
Casey Bodley [Thu, 11 Jun 2026 18:42:57 +0000 (14:42 -0400)]
Reapply "qa/rgw/crypt: disable failing kmip testing"

This reverts commit fd2046798198db20e45b8abbb3cc866c9967fb88.

kmip tests are failing again on ubuntu 24 because PyKMIP doesn't support
python 3.12. we'll be removing ubuntu 22 from main, so can't just pin the
test to that distro in the meantime

we're expecing the nvmeof team to add python 3.12 support in our
ceph/PyKMIP fork, and can reenable kmip testing once that happens

Fixes: https://tracker.ceph.com/issues/76995
Signed-off-by: Casey Bodley <cbodley@redhat.com>
(cherry picked from commit d27261d0c2641e92365cf590684567c58bbfb905)
Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
5 weeks agoqa/rgw: Remove 'force-branch' from s3tests configs 69716/head
Adam C. Emerson [Wed, 17 Jun 2026 18:33:45 +0000 (14:33 -0400)]
qa/rgw: Remove 'force-branch' from s3tests configs

Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
(cherry picked from commit 5bc34f8ed84aa2f205ccfe04fff9e3ec5ced8727)
Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
5 weeks agoqa/rgw: Run s3-tests from within the Ceph repo
Adam C. Emerson [Mon, 29 Sep 2025 21:11:09 +0000 (17:11 -0400)]
qa/rgw: Run s3-tests from within the Ceph repo

Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
(cherry picked from commit d0c88cb53c6848e9bcd7c4ea90f21f1aa6912242)
Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
5 weeks agotest/rgw: Include s3-tests in Ceph repo
Adam C. Emerson [Thu, 11 Jun 2026 22:46:45 +0000 (18:46 -0400)]
test/rgw: Include s3-tests in Ceph repo

Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
(cherry picked from commit 868d20182a08685265e0b5831bad703fdcf3189b)
Signed-off-by: Adam C. Emerson <aemerson@redhat.com>
5 weeks agomgr/dashboard: skip the table when an nvmeof cli result has no columns
Kefu Chai [Tue, 23 Jun 2026 07:43:28 +0000 (15:43 +0800)]
mgr/dashboard: skip the table when an nvmeof cli result has no columns

The dashboard leaves prettytable unpinned.  prettytable commit 2574492 ("Apply
some Pylint rules (PLR)", #436) rewrote _stringify_row()'s row_height as
`max(_get_size(c)[1] for c in row)`, which raises ValueError("max() iterable
argument is empty") on a row with no cells.  The change is undocumented and
shipped in 3.18.0; get_string() trips on it when a table has a row but no
columns.

AnnotatedDataTextOutputFormatter builds such a table for an empty result, or
one whose only field is status or error_message, so NvmeofCLICommand.call()
returns -EINVAL and the command fails.  This broke run-tox-mgr-dashboard-py3
once the tox virtualenv picked up prettytable 3.18.0.

Return an empty string when there are no columns instead of formatting a
degenerate table.

Fixes: https://tracker.ceph.com/issues/77589
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
(cherry picked from commit 9dd7fd0b0a78871f7f71c89a061490f3ffc97605)

5 weeks agorgw: restore constant-time GCM tag comparison in ISA-L path 69683/head
Matthew N. Heler [Tue, 9 Jun 2026 02:13:50 +0000 (21:13 -0500)]
rgw: restore constant-time GCM tag comparison in ISA-L path

a8ed43bfc05 replaced ct_memeq with memcmp in the ISA-L GCM accelerator,
making tag verification and the key-cache compare non-constant-time.
Restore ct_memeq for both; the OpenSSL and EVP paths already compare in
constant time.

Signed-off-by: Matthew N. Heler <matthew.heler@hotmail.com>
(cherry picked from commit 3a81a126089e201bcc0ddfcd800705d43f7d3b45)

5 weeks agorgw: fix AES-256-GCM key/IV reuse on multipart part re-upload
Matthew Heler [Fri, 5 Jun 2026 15:48:32 +0000 (10:48 -0500)]
rgw: fix AES-256-GCM key/IV reuse on multipart part re-upload

Re-uploading the same part number in a GCM multipart upload encrypted the new
data under the same key and IV as the first upload, since the IV is
part_number||chunk_index and the part key came from the part number alone. GCM
requires a unique IV per key; reusing one to encrypt different data weakens its
confidentiality and integrity guarantees.

Generate a random 16-byte salt on each UploadPart and fold it into the part key,
HMAC(ObjectKey, BE32(part) || salt), so every upload gets a fresh key. The salt
rides RGWUploadPartInfo, and complete stores the selected part's salt in
RGW_ATTR_CRYPT_PART_NUMS, which now holds (part, salt) pairs. GET reads it back
to re-derive the key, and an empty salt reproduces the old derivation so unsalted
parts still decrypt.

Signed-off-by: Matthew N. Heler <matthew.heler@hotmail.com>
(cherry picked from commit 57e0250f5738dcea3adf4193d82f00b90c52c136)

6 weeks agorgw/restore: take the hash mod HASH_PRIME when picking a shard 69606/head
Matthew N. Heler [Tue, 9 Jun 2026 22:21:39 +0000 (17:21 -0500)]
rgw/restore: take the hash mod HASH_PRIME when picking a shard

choose_oid fed ceph_str_hash_linux straight into % max_objs, and the
low bits of that hash are weak enough that similar object names keep
landing on the same shard. LC takes the hash mod HASH_PRIME first for
exactly this reason, do the same here.

Signed-off-by: Matthew N. Heler <matthew.heler@hotmail.com>
(cherry picked from commit d2e487acfe9ed45b30d879844bfdbc7d5f1ff343)

6 weeks agomon: add monitor RocksDB backup and restore 69586/head
Matthew N. Heler [Mon, 18 May 2026 01:57:01 +0000 (20:57 -0500)]
mon: add monitor RocksDB backup and restore

Implements an opt-in backup mechanism for the monitor using
rocksdb::BackupEngine. Backups run on a schedule when
mon_backup_interval is set, or are triggered manually via
`ceph tell mon.* backup`. Cleanup keeps the last N, hourly,
and daily snapshots, with a free-space guard. Off by default.

Restore is offline: stop the mon and run
  ceph-mon --restore-backup <dir> --yes-i-really-mean-it
optionally with --backup-version (BackupEngine logical version,
as shown by --list-backups). The mon keyring is stashed alongside
the RocksDB backup so a wiped mon_data is recovered end-to-end,
and kv_backend is stamped back when missing.

Co-authored-by: Daniel Poelzleithner <poelzleithner@b1-systems.de>
Signed-off-by: Matthew N. Heler <matthew.heler@hotmail.com>
(cherry picked from commit 3a9ae41e2a8fd614d67e3dac39d28ddf5dd6ca4a)

6 weeks agomgr/DaemonServer: Aggregate and globally sort OSDs for ok-to-upgrade 69549/head
Sridhar Seshasayee [Wed, 6 May 2026 15:11:33 +0000 (20:41 +0530)]
mgr/DaemonServer: Aggregate and globally sort OSDs for ok-to-upgrade

The 'ok-to-upgrade' command output sorting did not scale accurately
when target CRUSH buckets contained multiple child buckets (e.g., a
chassis containing multiple hosts). OSDs were previously sorted
individually per child bucket and appended sequentially. This created
fragmented, per-host sort segments rather than a globally sorted list
for the parent bucket.

Changes:

1. Fix the issue above by aggregating all child OSDs into a single vector prior
to executing a single, global sort operation based on PG counts. Additionally,
optimize memory efficiency and future-proof the logic by reserving continuous
vector blocks to avoid dynamic heap reallocations.

2. Add integration tests with chassis and rack based CRUSH hierarchies which
verifies the ok-to-upgrade functionality. In addition, the tests crucially
verify the order of OSDs returned is according to the ascending order of
acting PG count. Additionally, make minor fix-ups to lines that determine the
length of a list in JSON response by removing the redundant "| bc".

Fixes: https://tracker.ceph.com/issues/77272
Signed-off-by: Sridhar Seshasayee <sridhar.seshasayee@ibm.com>
(cherry picked from commit 6761549c5a719f541a850f243318a1212b76c0a6)

7 weeks agoMerge PR #69407 into umbrella
Patrick Donnelly [Thu, 11 Jun 2026 14:29:30 +0000 (10:29 -0400)]
Merge PR #69407 into umbrella

* refs/pull/69407/head:
umbrella: doc: update release-checklist
.github/milestone: add umbrella

Reviewed-by: Yuri Weinstein <yweins@redhat.com>
7 weeks agoumbrella: doc: update release-checklist 69407/head
Patrick Donnelly [Thu, 11 Jun 2026 03:27:51 +0000 (23:27 -0400)]
umbrella: doc: update release-checklist

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks ago.github/milestone: add umbrella
Patrick Donnelly [Wed, 10 Jun 2026 22:25:16 +0000 (18:25 -0400)]
.github/milestone: add umbrella

Fixes: https://tracker.ceph.com/issues/77308
Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
(cherry picked from commit aca3a35a875bf8330180faf3967bf75c849d415c)

7 weeks agoMerge PR #69405 into umbrella
Patrick Donnelly [Thu, 11 Jun 2026 01:56:20 +0000 (21:56 -0400)]
Merge PR #69405 into umbrella

* refs/pull/69405/head:
umbrella: doc: add nightlies
umbrella: doc: add release name to redmine
umbrella: doc: setup release redirects
umbrella: doc: add releases links to toc
umbrella: doc/dev/release-checkslists: remove past release notes
umbrella: doc/dev/release-checklists: branch created

Reviewed-by: Yuri Weinstein <yweins@redhat.com>
7 weeks agoumbrella: doc: add nightlies 69405/head
Patrick Donnelly [Wed, 10 Jun 2026 22:33:18 +0000 (18:33 -0400)]
umbrella: doc: add nightlies

This will be tracked via: https://tracker.ceph.com/issues/77309

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoumbrella: doc: add release name to redmine
Patrick Donnelly [Wed, 10 Jun 2026 22:24:26 +0000 (18:24 -0400)]
umbrella: doc: add release name to redmine

Already done.

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoumbrella: doc: setup release redirects
Patrick Donnelly [Wed, 10 Jun 2026 22:23:15 +0000 (18:23 -0400)]
umbrella: doc: setup release redirects

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoumbrella: doc: add releases links to toc
Patrick Donnelly [Fri, 18 Nov 2022 19:13:01 +0000 (14:13 -0500)]
umbrella: doc: add releases links to toc

Signed-off-by: Patrick Donnelly <pdonnell@redhat.com>
(cherry picked from commit 8cf9ad62949516666ad0f2c0bb7726ef68e4d666)
Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
Conflicts:
doc/index.rst: index changes

7 weeks agoumbrella: doc/dev/release-checkslists: remove past release notes
Patrick Donnelly [Wed, 10 Jun 2026 22:19:25 +0000 (18:19 -0400)]
umbrella: doc/dev/release-checkslists: remove past release notes

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoumbrella: doc/dev/release-checklists: branch created
Patrick Donnelly [Wed, 10 Jun 2026 22:18:58 +0000 (18:18 -0400)]
umbrella: doc/dev/release-checklists: branch created

Signed-off-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoMerge PR #66726 into main v21.0.1
Patrick Donnelly [Wed, 10 Jun 2026 18:30:59 +0000 (14:30 -0400)]
Merge PR #66726 into main

* refs/pull/66726/head:
doc: Update documentation to reflect new functionality
test: Add integration tests for EC Omap operations and recovery
osd: Hook up omap operations in EC pools
osd: Allow for recovery of OMAP header and entries in EC pools
doc: Write design document to explain the reasoning behind implementing this feature
osd: Introduce functions required for EC OMAP support
osd: Add ECOmapJournal class and relocate OmapUpdateType enum class

Reviewed-by: Bill Scales <bill_scales@uk.ibm.com>
Reviewed-by: Alex Ainscow <aainscow@uk.ibm.com>
Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agoMerge pull request #69051 from mheler/wip-rgw-http-reqs-lock
mheler [Wed, 10 Jun 2026 18:11:19 +0000 (13:11 -0500)]
Merge pull request #69051 from mheler/wip-rgw-http-reqs-lock

rgw/http: take reqs_lock when appending to reqs_change_state

7 weeks agoMerge pull request #68784 from mheler/wip-checksum-special-char
mheler [Wed, 10 Jun 2026 18:10:51 +0000 (13:10 -0500)]
Merge pull request #68784 from mheler/wip-checksum-special-char

rgw/cloud-transition: url-encode rgwx-source-key metadata header

7 weeks agoMerge pull request #69256 from ronen-fr/wip-rf-stshards
Ronen Friedman [Wed, 10 Jun 2026 15:31:58 +0000 (18:31 +0300)]
Merge pull request #69256 from ronen-fr/wip-rf-stshards

crimson/osd: avoid calling get_sharded_store() for obj size

Reviewed-by: Kefu Chai <k.chai@proxmox.com>
Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
Reviewed-by: Matan Breizman <mbreizma@redhat.com>
7 weeks agoMerge pull request #68888 from MattyWilliams22/mw-peering-state-rollforward
Matty Williams [Wed, 10 Jun 2026 15:20:23 +0000 (16:20 +0100)]
Merge pull request #68888 from MattyWilliams22/mw-peering-state-rollforward

osd: Fix condition for rolling forward pg log entries

Reviewed-by: Alex Ainscow <aainscow@uk.ibm.com>
Reviewed-by: Bill Scales <bill_scales@uk.ibm.com>
7 weeks agoMerge pull request #69276 from afreen23/worktree-umbrella-release-notes
Afreen Misbah [Wed, 10 Jun 2026 14:33:03 +0000 (20:03 +0530)]
Merge pull request #69276 from afreen23/worktree-umbrella-release-notes

doc: add Dashboard and Monitoring release notes for Umbrella

Reviewed-by: Afreen Misbah <afreen@ibm.com>
Reviewed-by: Naman Munet <nmunet@redhat.com>
7 weeks agoMerge pull request #68368 from kginonredhat/issue-75389-yaml-and-jinja2-deps-on-cento...
David Galloway [Wed, 10 Jun 2026 14:32:33 +0000 (10:32 -0400)]
Merge pull request #68368 from kginonredhat/issue-75389-yaml-and-jinja2-deps-on-centos-distro

ceph.spec: declare PyYAML and Jinja2 Requires for cephadm RPM

7 weeks agodoc: add Dashboard and Monitoring release notes for Umbrella 69276/head
Afreen Misbah [Mon, 25 May 2026 23:10:46 +0000 (04:40 +0530)]
doc: add Dashboard and Monitoring release notes for Umbrella

Signed-off-by: Afreen Misbah <afreen23@gmail.com>
7 weeks agoMerge pull request #68984 from Jayaprakash-ibm/wip-faster-alloc-recovery-testing
Jaya Prakash [Wed, 10 Jun 2026 11:31:07 +0000 (17:01 +0530)]
Merge pull request #68984 from Jayaprakash-ibm/wip-faster-alloc-recovery-testing

qa: Add Teuthology tests for BlueStore faster allocation recovery

Reviewed-by: Jaya Prakash <jayaprakash@ibm.com>
7 weeks agoMerge pull request #64369 from aclamk/aclamk-bs-faster-start-more
Jaya Prakash [Wed, 10 Jun 2026 11:30:14 +0000 (17:00 +0530)]
Merge pull request #64369 from aclamk/aclamk-bs-faster-start-more

bluestore: Faster allocation recovery - evolution

Reviewed-by: Jaya Prakash <jayaprakash@ibm.com>
7 weeks agoMerge pull request #68981 from aclamk/aclamk-kv-divide-range
Jaya Prakash [Wed, 10 Jun 2026 11:28:10 +0000 (16:58 +0530)]
Merge pull request #68981 from aclamk/aclamk-kv-divide-range

kv/KeyValueDB: New utility function util_divide_key_range

Reviewed-by: Jaya Prakash <jayaprakash@ibm.com>
7 weeks agoMerge pull request #69364 from eameh-LF/wip-doc-77191
Ilya Dryomov [Wed, 10 Jun 2026 10:00:45 +0000 (12:00 +0200)]
Merge pull request #69364 from eameh-LF/wip-doc-77191

doc/man: Remove stale EOL release names from deprecation notices

Reviewed-by: Ilya Dryomov <idryomov@gmail.com>
Reviewed-by: Anthony D'Atri <anthony.datri@gmail.com>
7 weeks agocrimson/osd: move get_max_object_size() to store level 69256/head
Ronen Friedman [Wed, 3 Jun 2026 05:40:25 +0000 (05:40 +0000)]
crimson/osd: move get_max_object_size() to store level

is_offset_and_length_valid() called get_sharded_store() locally to
obtain the store-specific max_object_size. On alien cores (where
smp::count > store_shard_nums), the local store is inactive and the
call hits assert(shard_store.get_status() == true).

As the max object size is a store-specific property and not a
store-shard one, there is no reason to acquire the
store shard to obtain it. Instead -
a get_max_object_size() method is added to the Store interface.

Fixes: https://tracker.ceph.com/issues/76946
Signed-off-by: Ronen Friedman <rfriedma@redhat.com>
7 weeks agoMerge pull request #68990 from rhcs-dashboard/carbon-filter
Nizamudeen A [Wed, 10 Jun 2026 05:02:26 +0000 (10:32 +0530)]
Merge pull request #68990 from rhcs-dashboard/carbon-filter

mgr/dashboard: carbonize table filters

Reviewed-by: Nizamudeen A <nia@redhat.com>
Reviewed-by: Afreen Misbah <afreen@ibm.com>
Reviewed-by: Naman Munet <nmunet@redhat.com>
7 weeks agoMerge pull request #69374 from sunyuechi/wip-catch2-disconnected-guard
Kefu Chai [Wed, 10 Jun 2026 03:18:07 +0000 (11:18 +0800)]
Merge pull request #69374 from sunyuechi/wip-catch2-disconnected-guard

cmake: disable Catch2 tests when Catch2 is unavailable

Reviewed-by: Kefu Chai <k.chai@proxmox.com>
7 weeks agoMerge pull request #69120 from tchaikov/wip-crimson-fix-move-rctx
Kefu Chai [Wed, 10 Jun 2026 01:52:35 +0000 (09:52 +0800)]
Merge pull request #69120 from tchaikov/wip-crimson-fix-move-rctx

crimson/osd: give each split child its own PeeringCtx

Reviewed-by: Aishwarya Mathuria <amathuri@redhat.com>
7 weeks agocmake: disable Catch2 tests when Catch2 is unavailable 69374/head
Sun Yuechi [Wed, 10 Jun 2026 00:13:53 +0000 (08:13 +0800)]
cmake: disable Catch2 tests when Catch2 is unavailable

debhelper on noble passes -DFETCHCONTENT_FULLY_DISCONNECTED=ON, so CPM
cannot fetch Catch2 and silently skips it, leaving no Catch2 targets
behind and breaking the generate step. Fall back to WITH_CATCH2=OFF
with a warning instead.

Signed-off-by: Sun Yuechi <sunyuechi@iscas.ac.cn>
7 weeks agoMerge pull request #61256 from irq0/wip/rgw-kms-cache
Adam Emerson [Tue, 9 Jun 2026 20:22:35 +0000 (16:22 -0400)]
Merge pull request #61256 from irq0/wip/rgw-kms-cache

RGW SSE-KMS secrets cache

Reviewed-by: Adam Emerson <aemerson@redhat.com>
7 weeks agoMerge pull request #69085 from dheart-joe/wip-reconstruct-allocations
Adam Kupczyk [Tue, 9 Jun 2026 19:06:36 +0000 (21:06 +0200)]
Merge pull request #69085 from dheart-joe/wip-reconstruct-allocations

os/bluestore: fix reallocation and corruption when shared_blob key is missing/undecodable

7 weeks agoMerge pull request #68837 from NitzanMordhai/wip-nitzan-cephtool-singleton-bluestore...
Laura Flores [Tue, 9 Jun 2026 18:59:59 +0000 (13:59 -0500)]
Merge pull request #68837 from NitzanMordhai/wip-nitzan-cephtool-singleton-bluestore-evicting-unresponsive-client

qa: ignore evicted client warnings for singletone bluestore

Reviewed-by: Radosław Zarzyński <Radoslaw.Adam.Zarzynski@ibm.com>
Reviewed-by: Yuri Weinstein <yweinste@ibm.com>
7 weeks agoMerge pull request #68825 from phlogistonjohn/jjm-smb-ctl-tool-fe
John Mulligan [Tue, 9 Jun 2026 18:21:32 +0000 (14:21 -0400)]
Merge pull request #68825 from phlogistonjohn/jjm-smb-ctl-tool-fe

smb: add a smb remote control client tool frontend

Reviewed-by: Avan Thakkar <athakkar@redhat.com>
Reviewed-by: Anoop C S <anoopcs@cryptolab.net>
7 weeks agoMerge pull request #65275 from ifed01/wip-ifed-no-buffered-wal
Igor Fedotov [Tue, 9 Jun 2026 15:51:59 +0000 (18:51 +0300)]
Merge pull request #65275 from ifed01/wip-ifed-no-buffered-wal

os/bluestore: do not use buffered IO for BlueFS WAL.

Reviewed-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoMerge pull request #69211 from Matan-B/wip-matanb-seastore-conflict-counters
Matan Breizman [Tue, 9 Jun 2026 13:53:42 +0000 (16:53 +0300)]
Merge pull request #69211 from Matan-B/wip-matanb-seastore-conflict-counters

crimsn/os/seastore: separate reset accounting from transaction creation

Reviewed-by: Xuehan Xu <xuxuehan@qianxin.com>
7 weeks agoos/bluestore: prevent reallocation and corruption when shared_blob key is missing... 69085/head
dheart [Tue, 9 Jun 2026 13:27:14 +0000 (21:27 +0800)]
os/bluestore: prevent reallocation and corruption when shared_blob key is missing/undecodable

When the shared_blob key is missing or fails to decode,
it is necessary to scan the blob's pextents directly as the sole authoritative source
to verify allocated blocks and prevent double-allocation.

Signed-off-by: dheart <dheart_joe@163.com>
7 weeks agoMerge pull request #69233 from tchaikov/wip-rgw-posix-thread-last
Casey Bodley [Tue, 9 Jun 2026 13:16:15 +0000 (09:16 -0400)]
Merge pull request #69233 from tchaikov/wip-rgw-posix-thread-last

rgw/posix: start the Inotify thread last, after the rest is built

Reviewed-by: Casey Bodley <cbodley@redhat.com>
7 weeks agodoc/man: Remove stale EOL release names from deprecation notices 69364/head
Emmanuel Ameh [Tue, 9 Jun 2026 12:40:03 +0000 (13:40 +0100)]
doc/man: Remove stale EOL release names from deprecation notices

ceph.rst: "osd create" deprecation notice cited "the Luminous release"
(2017, EOL 2020). Update to a plain deprecation statement directing
users to the replacement command (osd new).

rbd.rst: cephx_require_signatures option deprecation cited "the Bobtail
release" (2013, EOL 2015) as context for why the option is deprecated.
Remove the EOL release name; retain the deprecation warning. Fix the
companion nocephx_require_signatures notice for consistency ("in a
future release" instead of "in the future").

Fixes: https://tracker.ceph.com/issues/77191
Signed-off-by: Emmanuel Ameh <eameh@contractor.linuxfoundation.org>
7 weeks agoMerge pull request #69253 from cbodley/wip-76725
Casey Bodley [Tue, 9 Jun 2026 12:24:19 +0000 (08:24 -0400)]
Merge pull request #69253 from cbodley/wip-76725

osdc: deliver neorados completions to associated executor

Reviewed-by: Adam Emerson <aemerson@redhat.com>
Reviewed-by: Shilpa Jagannath <smanjara@redhat.com>
7 weeks agoMerge pull request #69246 from eameh-LF/i77075
eameh-LF [Tue, 9 Jun 2026 12:06:30 +0000 (13:06 +0100)]
Merge pull request #69246 from eameh-LF/i77075

doc/cephadm: fix typo and missing quote in activate-existing-osds

7 weeks agoMerge pull request #65792 from aclamk/aclamk-bs-onode-stall-fix
Jaya Prakash [Tue, 9 Jun 2026 11:53:16 +0000 (17:23 +0530)]
Merge pull request #65792 from aclamk/aclamk-bs-onode-stall-fix

os/bluestore: Fix problem with onode cache causing stalls

Reviewed-by: Igor Fedotov <igor.fedotov@croit.io>
7 weeks agoMerge pull request #68798 from aclamk/aclamk-bs-fix-stray-spanning-blobs
Jaya Prakash [Tue, 9 Jun 2026 11:52:57 +0000 (17:22 +0530)]
Merge pull request #68798 from aclamk/aclamk-bs-fix-stray-spanning-blobs

os/bluestore: Fix ExtentMap::reshard produce stray spanning blobs

Reviewed-by: Igor Fedotov <igor.fedotov@croit.io>
7 weeks agodoc: Update documentation to reflect new functionality 66726/head
Matty Williams [Mon, 23 Feb 2026 16:32:13 +0000 (16:32 +0000)]
doc: Update documentation to reflect new functionality

https://tracker.ceph.com/issues/74188
Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agotest: Add integration tests for EC Omap operations and recovery
Matty Williams [Tue, 23 Dec 2025 13:42:37 +0000 (13:42 +0000)]
test: Add integration tests for EC Omap operations and recovery

Assisted-by: Bob
Used for writing tests following the pattern of existing tests.

Fixes: https://tracker.ceph.com/issues/74188
Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agoosd: Hook up omap operations in EC pools
Matty Williams [Mon, 18 May 2026 09:09:32 +0000 (10:09 +0100)]
osd: Hook up omap operations in EC pools

Add pool flag to determine if omap operations are supported in a pool.
- Currently disabled in EC pools (will later be enabled for Fast EC pools)
Require all osds to have umbrella or later release version to enable pool flag.
Change recovery reads to use journal updates.
Clear the journal for a new epoch.
Set omap_complete accurately before recovery.
Encode omap updates and add entry to journal.
Decode omap updates, apply updates to object store, then remove from journal.
Change omap reads in PrimaryLogPG to use PGBackend functions, including omap updates from journal.

Assisted-by: Bob
Used for debugging and copying patterns (e.g. implementing REPLACE type to match MODIFY).

Fixes: https://tracker.ceph.com/issues/74188
Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agoosd: Allow for recovery of OMAP header and entries in EC pools
Matty Williams [Tue, 12 May 2026 15:11:17 +0000 (16:11 +0100)]
osd: Allow for recovery of OMAP header and entries in EC pools

Add omap fields to read_request_t, read_result_t, ECSubRead and ECSubReadReply.
Read and write omap header and entries if !omap_complete.
Require omap_complete to finish recovery.

Fixes: https://tracker.ceph.com/issues/74244
Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agodoc: Write design document to explain the reasoning behind implementing this feature
Matty Williams [Tue, 24 Feb 2026 15:16:28 +0000 (15:16 +0000)]
doc: Write design document to explain the reasoning behind implementing this feature

Assisted-by: Bob
Used to create the first draft of the design document.

https://tracker.ceph.com/issues/74187
Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agoosd: Introduce functions required for EC OMAP support
Matty Williams [Fri, 12 Dec 2025 11:21:10 +0000 (11:21 +0000)]
osd: Introduce functions required for EC OMAP support

Introduced a "supports_omap" pool flag which is always enabled for Replicated pools and currently always disabled for EC pools.
Introduced wrappers around omap read operations in PGBackend to include updates from the journal in EC pools with optimisations enabled.
Introduced a function for encoding an EC_OMAP operation in the ObjectModDesc::Visitor class and a function for committing an operation in the Trimmer struct.

Signed-off-by: Matty Williams <Matty.Williams@ibm.com>
7 weeks agoMerge pull request #69033 from kchheda3/fix-76729-notif-eventtime-race
Yuval Lifshitz [Tue, 9 Jun 2026 07:58:15 +0000 (10:58 +0300)]
Merge pull request #69033 from kchheda3/fix-76729-notif-eventtime-race

rgw/notification: fix zero eventTime in bucket notifications on concurrent PUT race

7 weeks agoMerge PR #68413 into main
Venky Shankar [Tue, 9 Jun 2026 01:32:00 +0000 (07:02 +0530)]
Merge PR #68413 into main

* refs/pull/68413/head:
mds: fix shutdown hang when ephemeral pins active and max_mds is 0
mds: fix crash in hash_into_rank_bucket() when max_mds is 0

Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
Reviewed-by: Venky Shankar <vshankar@redhat.com>
7 weeks agoMerge pull request #69165 from sunyuechi/wip-addcephtest-catch2-imported-target
Kefu Chai [Mon, 8 Jun 2026 23:37:28 +0000 (07:37 +0800)]
Merge pull request #69165 from sunyuechi/wip-addcephtest-catch2-imported-target

cmake/AddCephTest: use namespaced Catch2 imported targets

Reviewed-by: Jesse F. Williamson <jfw@ibm.com>
7 weeks agoMerge PR #69337 into main
Patrick Donnelly [Mon, 8 Jun 2026 22:31:53 +0000 (18:31 -0400)]
Merge PR #69337 into main

* refs/pull/69337/head:
doc: governance/csc: update email address

Reviewed-by: Joseph Mundackal <jmundackal@bloomberg.net>
Reviewed-by: Anthony D Atri <anthony.datri@gmail.com>
Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
7 weeks agodoc: governance/csc: update email address 69337/head
Yehuda Sadeh Weinraub [Mon, 8 Jun 2026 18:38:26 +0000 (11:38 -0700)]
doc: governance/csc: update email address

yehuda@redhat.com -> yehuda@ui.com

Signed-off-by: Yehuda Sadeh Weinraub <yehuda@ui.com>
7 weeks agoMerge pull request #69176 from Ericmzhang/wip-fix-pg_autoscaler-tests
Ericmzhang [Mon, 8 Jun 2026 19:12:11 +0000 (12:12 -0700)]
Merge pull request #69176 from Ericmzhang/wip-fix-pg_autoscaler-tests

qa: Fix pg autoscaler tests

7 weeks agoMerge pull request #69315 from sunyuechi/wip-sccache-riscv64
Zack Cerza [Mon, 8 Jun 2026 18:37:07 +0000 (12:37 -0600)]
Merge pull request #69315 from sunyuechi/wip-sccache-riscv64

Dockerfile.build: bump sccache and fetch it on riscv64

7 weeks agoqa/suites: add faster allocation recovery thrashing suite 68984/head
Jaya Prakash [Mon, 18 May 2026 19:57:50 +0000 (19:57 +0000)]
qa/suites: add faster allocation recovery thrashing suite

Signed-off-by: Jaya Prakash <jayaprakash@ibm.com>
7 weeks agoqa/workunits: add EC fio workload for allocation recovery testing
Jaya Prakash [Mon, 18 May 2026 19:57:33 +0000 (19:57 +0000)]
qa/workunits: add EC fio workload for allocation recovery testing

Signed-off-by: Jaya Prakash <jayaprakash@ibm.com>
7 weeks agoos/bluestore: Add printout to CBT's recovery-compare command 64369/head
Adam Kupczyk [Fri, 29 May 2026 11:16:39 +0000 (11:16 +0000)]
os/bluestore: Add printout to CBT's recovery-compare command

1) recovery-compare prints on stdout
2) gracefully rejects comparing when multithreaded not enabled

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Add bluestore_debug_fast_recovery_compare_chance
Adam Kupczyk [Tue, 19 May 2026 19:36:37 +0000 (19:36 +0000)]
os/bluestore: Add bluestore_debug_fast_recovery_compare_chance

The setting is used for testing purposes only.
It allows to force compare if required,
or set chance to use in teuthology thrash tests.

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Make OnodeScan use just one Blob
Adam Kupczyk [Mon, 7 Jul 2025 10:16:43 +0000 (10:16 +0000)]
os/bluestore: Make OnodeScan use just one Blob

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Tell OnodeScan to skip decoding checksums
Adam Kupczyk [Mon, 7 Jul 2025 10:02:01 +0000 (10:02 +0000)]
os/bluestore: Tell OnodeScan to skip decoding checksums

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Adapt multithread recovery
Adam Kupczyk [Mon, 7 Jul 2025 07:24:42 +0000 (07:24 +0000)]
os/bluestore: Adapt multithread recovery

Adapt multithread recovery to modified ExtentDecoder interface.

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Multithreaded allocation recovery
Adam Kupczyk [Thu, 3 Jul 2025 08:04:01 +0000 (08:04 +0000)]
os/bluestore: Multithreaded allocation recovery

Added multithreading processing for allocation recovery.
Added new config "bluestore_allocation_recovery_threads".

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Add "recovery-compare" action to CBT
Adam Kupczyk [Tue, 1 Jul 2025 13:25:38 +0000 (13:25 +0000)]
os/bluestore: Add "recovery-compare" action to CBT

New command compares 2 recovery modes:
 - legacy
 - new multithreaded
The command is hidden - it does not show in help.
Its role is devel & test only.

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>
7 weeks agoos/bluestore: Add new onode recovery method
Adam Kupczyk [Tue, 1 Jul 2025 13:47:14 +0000 (13:47 +0000)]
os/bluestore: Add new onode recovery method

Added read_allocation_from_onodes_mt function
  (originally copied from read_allocation_from_onodes).
Added Decoder_AllocationsAndStatFS class
  (originally copied from ExtentDecoderpartial).

There are significant differences from originals:
- shared blobs are not scanned at all
- to not account allocations more than once,
  collisions are detected on SimpleBitmap level;
  only the first onode referencing shared blob will mark allocation
- Blobs are not preserved
- instead we remember only if blob or spanning blob was compressed

The underlying logic is make recovery faster and prepare for
multithread refactor.

Signed-off-by: Adam Kupczyk <akupczyk@ibm.com>