]> git-server-git.apps.pok.os.sepia.ceph.com Git - ceph.git/log
ceph.git
3 days agoqa: Add cephadm-signed mtls test in nvmeof/mtls_test.sh 70612/head
Vallari Agrawal [Thu, 9 Apr 2026 19:30:51 +0000 (01:00 +0530)]
qa: Add cephadm-signed mtls test in nvmeof/mtls_test.sh

Expand nvmeof mtls test to include cephadm-signed cert
(ssl=true + enable_auth=true, no certs)

Also improve wait_for_service() logic to assert all
gateways are running.

Fixes: https://tracker.ceph.com/issues/78295
Signed-off-by: Vallari Agrawal <vallari.agrawal@ibm.com>
(cherry picked from commit 672b422cc9a3b10207e0f5fc312358ab7fdb2e80)

4 days agoMerge PR #69765 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:21:10 +0000 (10:21 -0400)]
Merge PR #69765 into umbrella

* refs/pull/69765/head:
rbd-mirror: Remove old non-primary demoted image snapshots on the local cluster
rbd-mirror: prune obsolete primary mirror snapshots after relocation
rbd-mirror: fix missing initialization of Peer UUID

Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
4 days agoMerge PR #69873 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:19:47 +0000 (10:19 -0400)]
Merge PR #69873 into umbrella

* refs/pull/69873/head:
librbd: fix use-after-free releasing object map locks in deep copy

Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
4 days agoMerge PR #70013 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:09:37 +0000 (10:09 -0400)]
Merge PR #70013 into umbrella

* refs/pull/70013/head:
journal/ObjectPlayer: don't acquire locks in destructor

Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
4 days agoMerge PR #70132 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:09:11 +0000 (10:09 -0400)]
Merge PR #70132 into umbrella

* refs/pull/70132/head:
qa/suites/rbd/valgrind: pin to centos_9.stream instead of rpm_latest

Reviewed-by: Ramana Raja <rraja@redhat.com>
4 days agoMerge PR #70256 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:01:41 +0000 (10:01 -0400)]
Merge PR #70256 into umbrella

* refs/pull/70256/head:
pybind/rbd: handle non-existent groups properly in list_snaps()

Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
4 days agoMerge PR #70287 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 14:01:13 +0000 (10:01 -0400)]
Merge PR #70287 into umbrella

* refs/pull/70287/head:
doc/rbd: elaborate on key-ref syntax (existing in S3Stream and new in NativeFormat)
qa/suites/rbd: add mon_host + key[-ref] coverage to migration-external
qa/suites/rbd: use client.0 entity in migration-external tests
librbd/migration/NativeFormat: support specifying mon_host and key via spec

4 days agoMerge PR #69778 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:57:09 +0000 (09:57 -0400)]
Merge PR #69778 into umbrella

* refs/pull/69778/head:
mgr/dashboard: Fix for EC profile creation modal scrollbar

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #69952 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:56:22 +0000 (09:56 -0400)]
Merge PR #69952 into umbrella

* refs/pull/69952/head:
mgr/dashboard: rbd-mirroring - hide create/import token buttons for

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #69954 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:55:02 +0000 (09:55 -0400)]
Merge PR #69954 into umbrella

* refs/pull/69954/head:
mgr/dashboard: Fix username validation for special characters by URL-encoding user lookup requests

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #69960 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:54:07 +0000 (09:54 -0400)]
Merge PR #69960 into umbrella

* refs/pull/69960/head:
mgr/dashboard: Fix daemon_name for NVMeoFClient

Reviewed-by: Afreen Misbah <afreen@ibm.com>
Reviewed-by: Redouane Kachach <rkachach@redhat.com>
4 days agoMerge PR #69982 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:53:43 +0000 (09:53 -0400)]
Merge PR #69982 into umbrella

* refs/pull/69982/head:
mgr/dashboard: fix-user-creation-validation

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70064 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:53:24 +0000 (09:53 -0400)]
Merge PR #70064 into umbrella

* refs/pull/70064/head:
mgr/dashboard: add account name duplicate check

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70070 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:53:00 +0000 (09:53 -0400)]
Merge PR #70070 into umbrella

* refs/pull/70070/head:
mgr/dashboard: teardown http requests for feature-toggle
mgr/dashboard: bump angular to 19.2.25

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70113 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:52:36 +0000 (09:52 -0400)]
Merge PR #70113 into umbrella

* refs/pull/70113/head:
mgr/dashboard: Align RGW Account Role forms with updated UX design

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70173 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:52:13 +0000 (09:52 -0400)]
Merge PR #70173 into umbrella

* refs/pull/70173/head:
mgr/dashboard: fix unncessary traceback when bucket not exist

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70176 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:51:53 +0000 (09:51 -0400)]
Merge PR #70176 into umbrella

* refs/pull/70176/head:
mgr/dashboard: fixing manual deployment option for osd creation flow
mgr/dashboard: fix cancel button
mgr/dashboard: Converting Create OSDs tearsheet type from Full to Wide
mgr/dashboard: carbonized OSD form component

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70194 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:51:24 +0000 (09:51 -0400)]
Merge PR #70194 into umbrella

* refs/pull/70194/head:
mgr/dashboard: Stale RGW user/account metadata cache in mgr after delete and recreate

Reviewed-by: Afreen Misbah <afreen@ibm.com>
4 days agoMerge PR #70270 into umbrella
Patrick Donnelly [Mon, 27 Jul 2026 13:50:38 +0000 (09:50 -0400)]
Merge PR #70270 into umbrella

* refs/pull/70270/head:
monitoring: Rename Ceph Object - Sync Overview - Replication (objects) from Source Zone

Reviewed-by: Afreen Misbah <afreen@ibm.com>
8 days agoMerge PR #70507 into umbrella
Patrick Donnelly [Fri, 24 Jul 2026 01:08:06 +0000 (21:08 -0400)]
Merge PR #70507 into umbrella

* refs/pull/70507/head:
qa/tests: drop parens from all health-code ignorelist entries
qa/tests: fix POOL_FULL ignorelist pattern in upgrade tests

Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
8 days agoqa/tests: drop parens from all health-code ignorelist entries 70507/head
Yuri Weinstein [Thu, 23 Jul 2026 19:32:51 +0000 (12:32 -0700)]
qa/tests: drop parens from all health-code ignorelist entries

Monitor.cc logs a "Health check cleared: <CODE> (was: ...)" line
for any health check clearing mid-test, with the code bare (no
parens) -- while the "Health check failed"/"Health check failed
(unmute)" lines wrap the code in parens. Every \(CODE\) entry in
this file therefore only matches the raise/failed form and misses
the corresponding cleared form, the same gap just fixed for
POOL_FULL.

Drop the escaped parens from the remaining entries so each matches
both log forms, consistent with the bare entries already present
(OSD_ROOT_DOWN, MDS_INSUFFICIENT_STANDBY, POOL_FULL).

Non-code entries (reached quota, overall HEALTH_, slow request,
noscrub, nodeep-scrub, osds down, insufficient standby, etc.) are
unchanged.

Fixes: https://tracker.ceph.com/issues/78149
Signed-off-by: Yuri Weinstein <yweinste@redhat.com>
(cherry picked from commit 8b261cd505a059b08f4187dc4ceedbf972d986ca)

8 days agoqa/tests: fix POOL_FULL ignorelist pattern in upgrade tests
Yuri Weinstein [Fri, 10 Jul 2026 15:53:36 +0000 (08:53 -0700)]
qa/tests: fix POOL_FULL ignorelist pattern in upgrade tests

\(POOL_FULL\) only matches the "Health check failed/update" cluster-log
format (code wrapped in parens). It misses the "Health check cleared"
format, where the code is not parenthesized (see src/mon/Monitor.cc),
so a pool clearing its full state mid-upgrade-test was not ignorelisted.
Drop the escaped parens so the pattern matches both forms, consistent
with other bare entries already in this file (OSD_ROOT_DOWN,
MDS_INSUFFICIENT_STANDBY).

Fixes: https://tracker.ceph.com/issues/78149
Signed-off-by: Yuri Weinstein <yweinste@redhat.com>
(cherry picked from commit c9ac93fdbbd5c8401e43cf297b43b8a85d2bd1bc)

11 days agoMerge PR #70297 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:56:28 +0000 (14:56 -0400)]
Merge PR #70297 into umbrella

* refs/pull/70297/head:
container: default FROM_IMAGE to Rocky Linux 10

Reviewed-by: Dan Mick <dmick@redhat.com>
Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
11 days agoMerge PR #70081 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:49:32 +0000 (14:49 -0400)]
Merge PR #70081 into umbrella

* refs/pull/70081/head:
doc/rbd: clarify mirror resync snapshot behavior

Reviewed-by: Anthony D'Atri <anthony.datri@gmail.com>
Reviewed-by: Ilya Dryomov <idryomov@redhat.com>
11 days agoMerge PR #70005 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:40:53 +0000 (14:40 -0400)]
Merge PR #70005 into umbrella

* refs/pull/70005/head:
qa/dnsmasq: remove unused backup/replace_resolv()
qa/dnsmasq: use managed dnsmasq instead of editing resolv.conf

11 days agoMerge PR #69817 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:40:08 +0000 (14:40 -0400)]
Merge PR #69817 into umbrella

* refs/pull/69817/head:
mgr: Properly set description in labeled get_perf_schema_python

Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
11 days agoMerge PR #69928 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:39:32 +0000 (14:39 -0400)]
Merge PR #69928 into umbrella

* refs/pull/69928/head:
mgr: filter root logger fallback

Reviewed-by: Kefu Chai <k.chai@proxmox.com>
11 days agoMerge PR #69941 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:38:53 +0000 (14:38 -0400)]
Merge PR #69941 into umbrella

* refs/pull/69941/head:
mgr: fix dispatch throttle bottleneck causing OSD connection timeouts

Reviewed-by: Radoslaw Zarzynski <rzarzyns@redhat.com>
11 days agoMerge PR #69962 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:38:31 +0000 (14:38 -0400)]
Merge PR #69962 into umbrella

* refs/pull/69962/head:
osd/PeeringState: add perf counters for PG rebuild times

Reviewed-by: Ronen Friedman <rfriedma@redhat.com>
11 days agoMerge PR #70029 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:36:50 +0000 (14:36 -0400)]
Merge PR #70029 into umbrella

* refs/pull/70029/head:
qa/test_mirroring: Add tests for mirroring checkpoints
doc: Add docs and pending release notes for mirroring checkpoints
mgr/mirroring,tools/cephfs_mirror: Handle checkpoint state transition
tools/cephfs_mirror: update checkpoint status during snapshot sync
tools/cephfs_mirror: Helper functions for mirroring checkpoints
mgr/mirroring: Add mirroring checkpoint CLIs
PendingReleaseNotes: add note for mutability of CephFS snapshot metadata
test_cephfs.py: add tests for do_snap_md_op()
pybind/cephfs: add python binding for ceph_do_snap_md_op()
tests/libcephfs: add tests for snap metadata mutations
libcephfs: provide API to mutate snapshot metadata
mds: add MDS code to allow snap metadata mutations
mds: rename a finisher method in Server.cc since...
mds: enable logging for snap.cc
client: add client code to allow snap metadata mutations
qa: Add mgr snapshot mirror status tests
doc/cephfs: Document fs snapshot mirror status mgr command
mgr/mirroring: make snapshot mirror metrics cache optional
mgr/mirroring: make snapshot mirror metrics cache TTL configurable
mgr/mirroring: handle missing cephfs_mirror object in metrics status
mgr/mirroring: Show default stats on newly added dir_root
mgr/mirroring: detect stale snapshot mirror sync metrics in omap
mgr/mirroring: cache fs snapshot mirror status omap metrics
mgr/mirroring: Reorder and format metrics output
mgr/mirroring: Add new interface to expose mirroring metrics
tools/cephfs_mirror: Remove persisted dir stats
tools/cephfs_mirror: Load persisted mirror metrics from omap
tools/cephfs_mirror: Persist metrics to omap
tools/cephfs_mirror: Adds capability to persist metrics
qa: Add tests for cephfs_mirror_directory perf counters
doc/cephfs: Document cephfs_mirror_directory counter dump metrics
cephfs_mirror: update per-directory last sync and summary perf counters
tools/cephfs_mirror: Expose per-directory snap metrics via perf counters
tools/cephfs_mirror: Add per-peer tick thread with configurable interval
tools/cephfs_mirror: Ignore duplicate directory acquire notifications
qa/cephfs: Add test for duplicate directory acquire notify
qa: Add mirror metrics testcases
doc: Update the mirroring doc with new metrics fields
qa: Fix the mirroring tests with new nested peer_status output
tools/cephfs_mirror: Nest peer_status metrics by dir path and peer uuid
tools/cephfs_mirror: Add datasync_queue_wait_duration metric
tools/cephfs_mirror: Add eta metrics
tools/cephfs_mirror: Add read/write throughput
tools/cephfs_mirror: Add crawl-state and sync-mode metric
tools/cephfs_mirror: Add inprogress bytes and files metric

Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
11 days agoMerge PR #69781 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:34:01 +0000 (14:34 -0400)]
Merge PR #69781 into umbrella

* refs/pull/69781/head:
mgr: handle SIGTERM/SIGINT in standby mgr to avoid CEPHADM_FAILED_DAEMON

Reviewed-by: Matan Breizman <mbreizma@redhat.com>
11 days agoMerge PR #69802 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:33:37 +0000 (14:33 -0400)]
Merge PR #69802 into umbrella

* refs/pull/69802/head:
mon/config: trim whitespace in config target

Reviewed-by: Laura Flores <lflores@redhat.com>
11 days agoMerge PR #69786 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:32:49 +0000 (14:32 -0400)]
Merge PR #69786 into umbrella

* refs/pull/69786/head:
rgw/lc: Warn against changing rgw_lc_max_objs once lifecycle is in use

11 days agoMerge PR #69553 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:28:42 +0000 (14:28 -0400)]
Merge PR #69553 into umbrella

* refs/pull/69553/head:
doc: add PendingReleaseNotes entry for rgw multisite DNS endpoint resolution
rgw/rest: add TODO for concurrent endpoint DNS resolution
rgw/multisite: fix endpoint unreachable detection in RGWRESTConn sync paths
rgw: store RGWEndpoint URL as boost::urls::url
doc/radosgw: expose rgw_rest_conn_connect_to_resolved_ips and rgw_rest_conn_ip_fail_timeout_secs
rgw: rename round-robin counters for brevity
rgw/zone: increase visibility into zone connections via admin socket
rgw: make CONN_STATUS_EXPIRE_SECS a cfg option
rgw/rest: track connection failures per-IP instead of per-endpoint
rgw/rest: remove unused headers
rgw/rest: consolidate endpoint_urls and resolved_endpoints into single vector
rgw: add operator<< for RGWEndpoint and simplify logging
rgw: track original URL within RGWEndpoint instead of separate member (refactor)
rgw: fix incomplete RGWRESTConn move constructor/assignment
rgw/rest: consolidate endpoint status tracking into ResolvedEndpoint
rgw/http: apply RGWEndpoint connect_to via libcurl CURLOPT_CONNECT_TO
rgw/rest: round-robin resolved endpoint IPs into curl CONNECT_TO mapping
rgw/rest: resolve multisite endpoints to all A/AAAA records (optional)
rgw: add rgw_resolve_endpoints_into_all_addresses config option
rgw/http: introduce RGWEndpoint to carry url + connect_to (refactor)

Reviewed-by: Adam C. Emerson <aemerson@redhat.com>
11 days agoMerge PR #69741 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:27:59 +0000 (14:27 -0400)]
Merge PR #69741 into umbrella

* refs/pull/69741/head:
tests/neorados: fix ceph_test_neorados_completions being not installed
common/async: async_cond::notify/cancel must post to handler's associated executor

Reviewed-by: Adam C. Emerson <aemerson@redhat.com>
11 days agoMerge PR #69723 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:26:49 +0000 (14:26 -0400)]
Merge PR #69723 into umbrella

* refs/pull/69723/head:
mgr/dashboard: Add "gw refresh_network" cmd
mgr/dashboard: bump nvmeof submodule to 1.8.2
mgr/dashboard: align nvmeof cli with missing parameters and functions from the old nvmeof cli
monitoring: fix NVMeoFMultipleNamespacesOfRBDImage

Reviewed-by: Vallari Agrawal <val.agl002@gmail.com>
Reviewed-by: Afreen Misbah <afreen@ibm.com>
11 days agoMerge PR #69731 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:26:05 +0000 (14:26 -0400)]
Merge PR #69731 into umbrella

* refs/pull/69731/head:
mgr/dashboard: fix daemon e2e

Reviewed-by: Afreen Misbah <afreen@ibm.com>
11 days agoMerge PR #69737 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:25:39 +0000 (14:25 -0400)]
Merge PR #69737 into umbrella

* refs/pull/69737/head:
rgw: use local error code in handle_individual_object()

Reviewed-by: Shilpa Jagannath <smanjara@redhat.com>
11 days agoMerge PR #69747 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:24:22 +0000 (14:24 -0400)]
Merge PR #69747 into umbrella

* refs/pull/69747/head:
rgw: reduce default thread pool size

Reviewed-by: Casey Bodley <cbodley@redhat.com>
11 days agoMerge PR #69452 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:17:51 +0000 (14:17 -0400)]
Merge PR #69452 into umbrella

* refs/pull/69452/head:
qa/workunits/mon: Update pg_autoscaler.sh in conjunction with https://github.com/ceph/ceph/pull/60492
src/common/options: Increase autoscaler PG target and overload values

Reviewed-by: Laura Flores <lflores@redhat.com>
11 days agoMerge PR #69580 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:13:13 +0000 (14:13 -0400)]
Merge PR #69580 into umbrella

* refs/pull/69580/head:
Containerfile: Support pulp repo URLs

Reviewed-by: Zack Cerza <zack@redhat.com>
11 days agoMerge PR #70267 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 18:11:35 +0000 (14:11 -0400)]
Merge PR #70267 into umbrella

* refs/pull/70267/head:
21.1.0
debian/rules: Also exclude librgw and radosgw from dwz
debian/rules: exclude ceph-osd-crimson from dwz compression
script/buildcontainer-setup: fix package install on debian trixie

Reviewed-by: Yuri Weinstein <yweins@redhat.com>
11 days agoMerge PR #70265 into umbrella
Patrick Donnelly [Mon, 20 Jul 2026 13:22:06 +0000 (09:22 -0400)]
Merge PR #70265 into umbrella

* refs/pull/70265/head:
debian/rules: Also exclude librgw and radosgw from dwz

Reviewed-by: Kefu Chai <k.chai@proxmox.com>
Reviewed-by: Patrick Donnelly <pdonnell@ibm.com>
13 days agodebian/rules: Also exclude librgw and radosgw from dwz 70265/head
David Galloway [Thu, 16 Jul 2026 20:25:07 +0000 (16:25 -0400)]
debian/rules: Also exclude librgw and radosgw from dwz

Fixes: https://tracker.ceph.com/issues/78313
Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit f17083b81e13822b7bc6535aad3188f8624dcaaa)

13 days agoMerge PR #70261 into umbrella
Patrick Donnelly [Sat, 18 Jul 2026 23:38:42 +0000 (19:38 -0400)]
Merge PR #70261 into umbrella

* refs/pull/70261/head:
debian/rules: exclude ceph-osd-crimson from dwz compression

Reviewed-by: Kefu Chai <k.chai@proxmox.com>
Reviewed-by: Yuri Weinstein <yweins@redhat.com>
2 weeks agocontainer: default FROM_IMAGE to Rocky Linux 10 70297/head
David Galloway [Fri, 17 Jul 2026 18:34:44 +0000 (14:34 -0400)]
container: default FROM_IMAGE to Rocky Linux 10

Since tentacle, the preferred base image for ceph containers has been
Rocky Linux 10, and the CI tag-naming logic in build.sh already assumes
rockylinux-10 is the default fromtag for every branch except reef and
squid.  The actual build default was never flipped, though: anything
that ran build.sh without FROM_IMAGE set (e.g. the release container
job in ceph-build) still got a CentOS Stream 9 base.

Flip the Containerfile ARG and the build.sh fallback to
docker.io/rockylinux/rockylinux:10 so umbrella and later build from
Rocky 10 by default.  Builds that want a different base can still pass
FROM_IMAGE explicitly, as the CI pipeline does.

Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit 194e58aa9be34a39310c29eb6504537b79b4b95a)

2 weeks agodoc/rbd: elaborate on key-ref syntax (existing in S3Stream and new in NativeFormat) 70287/head
Ilya Dryomov [Wed, 1 Jul 2026 11:17:40 +0000 (13:17 +0200)]
doc/rbd: elaborate on key-ref syntax (existing in S3Stream and new in NativeFormat)

Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
(cherry picked from commit c7fddec189e8c6eac08820b73ce44dbd795a28ae)

2 weeks agoqa/suites/rbd: add mon_host + key[-ref] coverage to migration-external
Ilya Dryomov [Mon, 29 Jun 2026 09:40:02 +0000 (11:40 +0200)]
qa/suites/rbd: add mon_host + key[-ref] coverage to migration-external

Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
(cherry picked from commit e246fad552854cdc942485047b6c9112524a9229)

2 weeks agoqa/suites/rbd: use client.0 entity in migration-external tests
Ilya Dryomov [Mon, 29 Jun 2026 11:39:50 +0000 (13:39 +0200)]
qa/suites/rbd: use client.0 entity in migration-external tests

Currently client.admin is passed for client_name and that doesn't
exercise client_name handling much as client.admin is the default.

When deploying multiple clusters the ceph task distributes keyrings
with only the key for the initial monitor and client.admin key.  Other
keys (e.g. client.0) are present only on their respective clusters, so
client.0's key for cluster2 needs to be obtained on cluster1 explicitly
with "ceph auth get".

Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
(cherry picked from commit 047bbbf863032c76532cc51eadd5368189a8b146)

2 weeks agolibrbd/migration/NativeFormat: support specifying mon_host and key via spec
Leonid Chernin [Mon, 2 Feb 2026 06:01:14 +0000 (08:01 +0200)]
librbd/migration/NativeFormat: support specifying mon_host and key via spec

migration:
           -take  secret_key from spec
           -take  mon_host from the spec
           -ignore source keyring, source ceph.conf - bypass them
           -added validations of the new way schema
           -get key from the KV DB if key in spec  is actually key-ref

Fixes: https://tracker.ceph.com/issues/68177
Signed-off-by: Leonid Chernin <leonidc@il.ibm.com>
Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
(cherry picked from commit 705d865ae77f4d3e80a9d3da3281ba6a824d7f29)

2 weeks agomonitoring: Rename Ceph Object - Sync Overview - Replication (objects) from Source... 70270/head
Aashish Sharma [Wed, 15 Jul 2026 08:21:23 +0000 (13:51 +0530)]
monitoring: Rename Ceph Object - Sync Overview - Replication (objects) from Source Zone

Replication (objects) from Source Zone title is misleading, the panel actually shows How many objects are being replicated per second

Fixes: https://tracker.ceph.com/issues/78230
Signed-off-by: Aashish Sharma <aasharma@redhat.com>
(cherry picked from commit ce24915123b4a44110da31ee760a928a0b1b71b9)

2 weeks ago21.1.0 70267/head v21.1.0
Ceph Release Team [Thu, 16 Jul 2026 23:19:28 +0000 (23:19 +0000)]
21.1.0

Signed-off-by: Ceph Release Team <ceph-maintainers@ceph.io>
2 weeks agodebian/rules: Also exclude librgw and radosgw from dwz
David Galloway [Thu, 16 Jul 2026 20:25:07 +0000 (16:25 -0400)]
debian/rules: Also exclude librgw and radosgw from dwz

Fixes: https://tracker.ceph.com/issues/78313
Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit f17083b81e13822b7bc6535aad3188f8624dcaaa)
(cherry picked from commit 789b62510d78b0aa7e42142ecb18033e560b70b8)
Signed-off-by: Yuri Weinstein <yweinste@redhat.com>
2 weeks agodebian/rules: exclude ceph-osd-crimson from dwz compression
Kefu Chai [Fri, 23 Jan 2026 07:04:47 +0000 (15:04 +0800)]
debian/rules: exclude ceph-osd-crimson from dwz compression

When building with DWZ enabled, the debian packaging fails with:
```
  dh_dwz: error: Aborting due to earlier error
```
Running the dwz command manually reveals the root cause:
```
  $ dwz -mdebian/ceph-osd-crimson/usr/lib/debug/.dwz/x86_64-linux-gnu/ceph-osd-crimson.debug \
    -M/usr/lib/debug/.dwz/x86_64-linux-gnu/ceph-osd-crimson.debug -- \
    debian/ceph-osd-crimson/usr/bin/ceph-osd-crimson \
    debian/ceph-osd-crimson/usr/bin/crimson-store-nbd
  dwz: debian/ceph-osd-crimson/usr/bin/ceph-osd-crimson: Too many DIEs, not optimizing
  dwz: Too few files for multifile optimization
```
The dwz tool has a limit on the number of DWARF DIEs (Debug Information
Entries) it can process. The ceph-osd-crimson binary, being a large C++
executable with extensive template usage, exceeds this limit, causing
dwz to exit with status 1 and fail the build.

This change excludes only ceph-osd-crimson from dwz processing using
the -X flag, allowing other binaries in the package to still benefit
from DWARF compression while avoiding the build failure.

Please note, we always disable DWZ in ceph-build's ceph-dev-pipeline.

Fixes: https://tracker.ceph.com/issues/78303
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
(cherry picked from commit 0b50da6588a464cb9d3758f93485055468f61549)
(cherry picked from commit 005656182db7871fe95c632b9a04a972b97c2cef)

2 weeks agodebian/rules: exclude ceph-osd-crimson from dwz compression 70261/head
Kefu Chai [Fri, 23 Jan 2026 07:04:47 +0000 (15:04 +0800)]
debian/rules: exclude ceph-osd-crimson from dwz compression

When building with DWZ enabled, the debian packaging fails with:
```
  dh_dwz: error: Aborting due to earlier error
```
Running the dwz command manually reveals the root cause:
```
  $ dwz -mdebian/ceph-osd-crimson/usr/lib/debug/.dwz/x86_64-linux-gnu/ceph-osd-crimson.debug \
    -M/usr/lib/debug/.dwz/x86_64-linux-gnu/ceph-osd-crimson.debug -- \
    debian/ceph-osd-crimson/usr/bin/ceph-osd-crimson \
    debian/ceph-osd-crimson/usr/bin/crimson-store-nbd
  dwz: debian/ceph-osd-crimson/usr/bin/ceph-osd-crimson: Too many DIEs, not optimizing
  dwz: Too few files for multifile optimization
```
The dwz tool has a limit on the number of DWARF DIEs (Debug Information
Entries) it can process. The ceph-osd-crimson binary, being a large C++
executable with extensive template usage, exceeds this limit, causing
dwz to exit with status 1 and fail the build.

This change excludes only ceph-osd-crimson from dwz processing using
the -X flag, allowing other binaries in the package to still benefit
from DWARF compression while avoiding the build failure.

Please note, we always disable DWZ in ceph-build's ceph-dev-pipeline.

Fixes: https://tracker.ceph.com/issues/78303
Signed-off-by: Kefu Chai <k.chai@proxmox.com>
(cherry picked from commit 0b50da6588a464cb9d3758f93485055468f61549)

2 weeks agopybind/rbd: handle non-existent groups properly in list_snaps() 70256/head
VinayBhaskar-V [Thu, 9 Jul 2026 16:00:26 +0000 (21:30 +0530)]
pybind/rbd: handle non-existent groups properly in list_snaps()

Calling `list_snaps()` on a non-existent group crashed the entire
Python process with a `free(): invalid pointer` error. This happened
because an unexpected `-ENOENT` error returned by `rbd_group_snap_list2`
and the exception was raised without updating the num_group_snaps to 0
from 10. Later the Cython `__dealloc__` wrapper attempted to free garbage
of uninitialized pointer spaces resulting in a crash.

Fix this behavior by resetting `self.num_group_snaps` to 0 inside
`GroupSnapIterator` before raising exception to prevent memory
corruption and accessing uninitialized pointers during cleanup call

Now `list_snaps()` consistently raise an `ObjectNotFound` error when
invoked on a non-existent group.
Also Extended the `TestGroups` test fixture with `self.dne_group` to
validate the expected error path for `list_snaps()`.

Fixes: https://tracker.ceph.com/issues/78034
Signed-off-by: VinayBhaskar-V <vvarada@redhat.com>
(cherry picked from commit 56bcf706cf91f0734899ae9de378c459a9b5af50)

2 weeks agoscript/buildcontainer-setup: fix package install on debian trixie
David Galloway [Fri, 10 Jul 2026 20:12:01 +0000 (16:12 -0400)]
script/buildcontainer-setup: fix package install on debian trixie

Debian removed software-properties-common from the archive in trixie,
so the flat apt-get install list fails there. The package was only
needed to provide add-apt-repository for llvm.sh, which installs
clang-19 from apt.llvm.org. Trixie ships clang-19 natively, so install
it from the distro instead; run-make.sh's prepare() then finds clang-19
and skips llvm.sh entirely. Ubuntu and older Debian releases still have
software-properties-common and keep the previous behavior.

Fixes: https://tracker.ceph.com/issues/78111
Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit b5697987ae86f51d7f16a1409ce9f2b542f3de52)
(cherry picked from commit 113c413d9be71091f7315861357a9024d1a6758e)

2 weeks agomgr/dashboard: Stale RGW user/account metadata cache in mgr after delete and recreate 70194/head
Naman Munet [Mon, 6 Jul 2026 17:15:22 +0000 (22:45 +0530)]
mgr/dashboard: Stale RGW user/account metadata cache in mgr after delete and recreate

fixes: https://tracker.ceph.com/issues/77544

Signed-off-by: Naman Munet <naman.munet@ibm.com>
(cherry picked from commit 6bf1347dbc2734ccbc6dcc386470b9f9656d842e)

2 weeks agomgr/dashboard: fixing manual deployment option for osd creation flow 70176/head
Devika Babrekar [Wed, 8 Jul 2026 13:10:12 +0000 (18:40 +0530)]
mgr/dashboard: fixing manual deployment option for osd creation flow

Signed-off-by: Devika Babrekar <devika.babrekar@ibm.com>
(cherry picked from commit b9e272c469250d2f1bcc4301117ca790ae1c4748)

2 weeks agomgr/dashboard: fix cancel button
Afreen Misbah [Tue, 7 Jul 2026 12:58:32 +0000 (18:28 +0530)]
mgr/dashboard: fix cancel button

Signed-off-by: Afreen Misbah <afreen@ibm.com>
(cherry picked from commit b35999b6435cf4bf10ebdfa0204621836d76bcf0)

2 weeks agomgr/dashboard: Converting Create OSDs tearsheet type from Full to Wide
Devika Babrekar [Thu, 2 Jul 2026 10:46:46 +0000 (16:16 +0530)]
mgr/dashboard: Converting Create OSDs tearsheet type from Full to Wide
Fixes: https://tracker.ceph.com/issues/77132
Signed-off-by: Devika Babrekar <devika.babrekar@ibm.com>
(cherry picked from commit 4eb4e25e5e6f14ed70fee5c41b2669cc0c557e19)

2 weeks agomgr/dashboard: carbonized OSD form component
Syed Ali Ul Hasan [Wed, 10 Jun 2026 17:30:36 +0000 (23:00 +0530)]
mgr/dashboard: carbonized OSD form component

Fixes: https://tracker.ceph.com/issues/68265
Signed-off-by: Syed Ali Ul Hasan <syedaliulhasan19@gmail.com>
(cherry picked from commit 78dad340ab180532082c39a52efc0dfce244ec6e)

2 weeks agomgr/dashboard: fix unncessary traceback when bucket not exist 70173/head
Nizamudeen A [Wed, 8 Jul 2026 09:27:33 +0000 (14:57 +0530)]
mgr/dashboard: fix unncessary traceback when bucket not exist

UI has an async validator which calls the GET bucket API to make sure
the bucket name doesn't exist, but the proxy
was not properly handling the http_status_codes which results in raising
a massive traceback in logs whenever you type things in the bucket name
field. So handling that gracefully by capturing the proper status codes
for both RequestException and DashboardException

BEFORE
```
File "/usr/share/ceph/mgr/dashboard/services/exception.py", line 47, in dashboard_exception_handler
return handler(*args, **kwargs)
File "/lib/python3.9/site-packages/cherrypy/_cpdispatch.py", line 54, in _call_
return self.callable(*self.args, **self.kwargs)
File "/usr/share/ceph/mgr/dashboard/controllers/_base_controller.py", line 263, in inner
ret = func(*args, **kwargs)
File "/usr/share/ceph/mgr/dashboard/controllers/_rest_controller.py", line 193, in wrapper
return func(*vpath, **params)
File "/usr/share/ceph/mgr/dashboard/controllers/rgw.py", line 357, in get
result = self.proxy(daemon_name, 'GET', 'bucket', {'bucket': bucket})
File "/usr/share/ceph/mgr/dashboard/controllers/rgw.py", line 213, in proxy
raise DashboardException(e, http_status_code=http_status_code, component='rgw')
dashboard.exceptions.DashboardException: RGW REST API failed request with status code 404
(b'{"Code":"NoSuchBucket","Message":"","RequestId":"tx00000f14e08c1af0d5615-006'
b'71f54a7-3bc6-default","HostId":"3bc6-default-default"}')
2024-10-28T09:08:55.990+0000 7f89e045b640  0 [dashboard INFO request] [::ffff:10.74.18.122:51853] [GET] [500] [0.007s] [admin] [200.0B] /api/rgw/bucket/bucket-das
```

AFTER
```
Jul 08 09:28:28 ceph-node-00 ceph-mgr[2243]: [dashboard ERROR dashboard.rest_client] RGW REST API failed GET req status: 404
Jul 08 09:28:28 ceph-node-00 ceph-mgr[2243]: [dashboard INFO dashboard.services.exception] Dashboard Exception: RGW REST API failed request with status code 404
                                             (b'{"Code":"NoSuchBucket","Message":"","RequestId":"tx00000c029a81e5d154844-006'
                                              b'a4e183c-14251-default","HostId":"14251-default-default"}')
Jul 08 09:28:28 ceph-node-00 ceph-mgr[2243]: [dashboard INFO dashboard.tools] [::ffff:192.168.100.1:60408] [GET] [404] [0.078s] [admin] [200.0B] /api/rgw/bucket/testsa
```

Fixes: https://tracker.ceph.com/issues/78038
Signed-off-by: Nizamudeen A <nia@redhat.com>
(cherry picked from commit 0ad33c0965dcd1ea325a70253c5ec8f8e2219936)

2 weeks agoMerge PR #70128 into umbrella
Patrick Donnelly [Mon, 13 Jul 2026 18:48:38 +0000 (14:48 -0400)]
Merge PR #70128 into umbrella

* refs/pull/70128/head:
script/buildcontainer-setup: fix package install on debian trixie

Reviewed-by: John Mulligan <jmulligan@redhat.com>
Reviewed-by: Casey Bodley <cbodley@redhat.com>
2 weeks agoqa/suites/rbd/valgrind: pin to centos_9.stream instead of rpm_latest 70132/head
Ilya Dryomov [Tue, 7 Jul 2026 09:29:44 +0000 (11:29 +0200)]
qa/suites/rbd/valgrind: pin to centos_9.stream instead of rpm_latest

This used to be the case before commit d4b977afdc59 ("qa/distros:
rename centos_latest.yaml to rpm_latest.yaml") and subsequent changes
to enable Rocky 10.  When paired with valgrind, python_api_tests* jobs
fail persistently on Ubuntu and Rocky, see [1].  Switch back until that
is resolved (or for as long as centos_9.stream remains in the mix).

[1] https://tracker.ceph.com/issues/74864

Fixes: https://tracker.ceph.com/issues/77982
Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
(cherry picked from commit 75c4e2897f03240d64efca1d6b95abfd8b799dac)

3 weeks agoscript/buildcontainer-setup: fix package install on debian trixie 70128/head
David Galloway [Fri, 10 Jul 2026 20:12:01 +0000 (16:12 -0400)]
script/buildcontainer-setup: fix package install on debian trixie

Debian removed software-properties-common from the archive in trixie,
so the flat apt-get install list fails there. The package was only
needed to provide add-apt-repository for llvm.sh, which installs
clang-19 from apt.llvm.org. Trixie ships clang-19 natively, so install
it from the distro instead; run-make.sh's prepare() then finds clang-19
and skips llvm.sh entirely. Ubuntu and older Debian releases still have
software-properties-common and keep the previous behavior.

Fixes: https://tracker.ceph.com/issues/78111
Signed-off-by: David Galloway <david.galloway@ibm.com>
(cherry picked from commit b5697987ae86f51d7f16a1409ce9f2b542f3de52)

3 weeks agomgr/dashboard: Align RGW Account Role forms with updated UX design 70113/head
Sagar Gopale [Thu, 25 Jun 2026 07:07:25 +0000 (12:37 +0530)]
mgr/dashboard: Align RGW Account Role forms with updated UX design

Signed-off-by: Sagar Gopale <sagar.gopale@ibm.com>
(cherry picked from commit 5f5e4093ad3dbe306593999a6b1ee0410f079af3)

3 weeks agodoc/rbd: clarify mirror resync snapshot behavior 70081/head
Super User [Wed, 8 Jul 2026 12:05:16 +0000 (17:35 +0530)]
doc/rbd: clarify mirror resync snapshot behavior

Add documentation explaining that rbd mirror resync only copies
image state and data up to the last mirror-snapshot for
snapshot-based mirroring. Document the requirement to create
a new mirror-snapshot on current primary so that new changes
(including resize operation) gets reflected on secondary after
resync.

Fixes: https://tracker.ceph.com/issues/78050
Signed-off-by: Miki Patel <miki.patel132@gmail.com>
(cherry picked from commit d80a038a00fa6e405a3df2ec66f7fdd303a4370c)

3 weeks agomgr/dashboard: teardown http requests for feature-toggle 70070/head
Nizamudeen A [Wed, 8 Jul 2026 05:21:50 +0000 (10:51 +0530)]
mgr/dashboard: teardown http requests for feature-toggle

angular 19.2.latest is aggressively tearing down the test env which
forces the active rxjs timer to emit one execution as it collapses which
produces a ghost http in the case of request wrapped in the timer
service. so for now flushing all the requests manually to prevent it

Fixes: https://tracker.ceph.com/issues/77729
Signed-off-by: Nizamudeen A <nia@redhat.com>
(cherry picked from commit fcbbca8949a97aa73bf59f318acd449f103b415c)

3 weeks agomgr/dashboard: bump angular to 19.2.25
Nizamudeen A [Fri, 26 Jun 2026 09:32:08 +0000 (15:02 +0530)]
mgr/dashboard: bump angular to 19.2.25

Fixes: https://tracker.ceph.com/issues/77729
Signed-off-by: Nizamudeen A <nia@redhat.com>
(cherry picked from commit 13383ddfa2592ac6bb5150dca691b5ccc5f4c0d8)

3 weeks agomgr/dashboard: add account name duplicate check 70064/head
Naman Munet [Tue, 30 Jun 2026 13:03:25 +0000 (18:33 +0530)]
mgr/dashboard: add account name duplicate check

fixes: https://tracker.ceph.com/issues/77828

Signed-off-by: Naman Munet <naman.munet@ibm.com>
(cherry picked from commit 8dafeeb22e5c7d2ea6cc4e52c91ac780ae9194fe)

3 weeks agoqa/test_mirroring: Add tests for mirroring checkpoints 70029/head
Karthik U S [Sat, 16 May 2026 03:15:50 +0000 (08:45 +0530)]
qa/test_mirroring: Add tests for mirroring checkpoints

Adding integration tests for validating the cephfs mirroring
checkpoints feature

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit fe7fa7b0761720181cf8735ac7b5f17080c1d8a7)

3 weeks agodoc: Add docs and pending release notes for mirroring checkpoints
Karthik U S [Thu, 4 Jun 2026 11:47:05 +0000 (17:17 +0530)]
doc: Add docs and pending release notes for mirroring checkpoints

Adding documentations and pending release notes for the mirroring
checkpoints feature.

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit 12bd34d69ac7e136d14d25ea274e99727b559870)

3 weeks agomgr/mirroring,tools/cephfs_mirror: Handle checkpoint state transition
Karthik U S [Wed, 10 Jun 2026 05:48:25 +0000 (11:18 +0530)]
mgr/mirroring,tools/cephfs_mirror: Handle checkpoint state transition

When a new checkpoint is being added or when the daemon gets restarted,
it will check whether the newly created checkpoint or any other old
checkpoints have already been mirrored onto the remote peer. If so, it
will transition to the correct state by checking for the highest snap id
present on the remote, and setting all the checkpoints which have snap id
lesser than or equal to that of the remote to COMPLETE. This is done by
sending an acquire notification from the mirroring module to the mirror
daemon, which is handled in the add_directory path, by adding the directory
to be checked for the state transition in the tick thread.
This path gets triggered:
a) when the daemon gets restarted
b) when the peer mapping changes
c) by sending the acquire notification from checkpoint add/now CLIs.
d) when mirroring module restarts

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit c64fe9636514d9e9b1c286a0b21b84569e266067)

3 weeks agotools/cephfs_mirror: update checkpoint status during snapshot sync
Karthik U S [Sat, 16 May 2026 00:50:03 +0000 (06:20 +0530)]
tools/cephfs_mirror: update checkpoint status during snapshot sync

When the sync completes or fails for a checkpointed snapshot, transition
the status of that checkpoint from CREATED/FAILED to COMPLETE/FAILED in
the do_sync_snaps() along with the timestamp of the event.

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit 37163436cc87a2be691202bb565c7b5007df026f)

3 weeks agotools/cephfs_mirror: Helper functions for mirroring checkpoints
Karthik U S [Tue, 12 May 2026 23:21:53 +0000 (04:51 +0530)]
tools/cephfs_mirror: Helper functions for mirroring checkpoints

Implementation of helper functions and data structures for the
snapshot based mirroring checkpoints feature.

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit 0a394ef438abd0248d078ea72e8ae5ce4f3d5e84)

3 weeks agomgr/mirroring: Add mirroring checkpoint CLIs
Karthik U S [Sat, 16 May 2026 00:49:59 +0000 (06:19 +0530)]
mgr/mirroring: Add mirroring checkpoint CLIs

Add mgr checkpoint add/remove/ls/now commands that read and write
checkpoint metadata on the primary filesystem via do_snap_md_op.

Sample CLIs:
ceph fs snapshot mirror checkpoint add <vol-name> <dir-root> <snap-name>
ceph fs snapshot mirror checkpoint remove <vol-name> <dir-root> <snap-name>
ceph fs snapshot mirror checkpoint ls <vol-name> <dir-root>
ceph fs snapshot mirror checkpoint now <vol-name> <dir-root>

Fixes: https://tracker.ceph.com/issues/73454
Signed-off-by: Karthik U S <karthik.u.s1@ibm.com>
(cherry picked from commit 959bc2eb07cacf70a4da09170e1c9bbed4b9423f)

3 weeks agoPendingReleaseNotes: add note for mutability of CephFS snapshot metadata
Rishabh Dave [Mon, 15 Jun 2026 11:09:25 +0000 (16:39 +0530)]
PendingReleaseNotes: add note for mutability of CephFS snapshot metadata

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 7f1f0e9eb2fabb88a8e44525cb8490a44a95a79d)

3 weeks agotest_cephfs.py: add tests for do_snap_md_op()
Rishabh Dave [Wed, 1 Apr 2026 12:32:28 +0000 (18:02 +0530)]
test_cephfs.py: add tests for do_snap_md_op()

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit df7abb7a7b7be952e30e2ac286b108c3cab5ca90)

3 weeks agopybind/cephfs: add python binding for ceph_do_snap_md_op()
Rishabh Dave [Wed, 1 Apr 2026 12:31:56 +0000 (18:01 +0530)]
pybind/cephfs: add python binding for ceph_do_snap_md_op()

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 8d5cf70ba6affe5fcb921509befc59c990c5a3bf)

3 weeks agotests/libcephfs: add tests for snap metadata mutations
Rishabh Dave [Wed, 19 Nov 2025 17:12:35 +0000 (22:42 +0530)]
tests/libcephfs: add tests for snap metadata mutations

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 446bd6af1123884912ffbb743950c99c6b699267)

3 weeks agolibcephfs: provide API to mutate snapshot metadata
Rishabh Dave [Mon, 17 Nov 2025 12:23:56 +0000 (17:53 +0530)]
libcephfs: provide API to mutate snapshot metadata

Fixes: https://tracker.ceph.com/issues/66293
Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 5726edbcf37ed9c3cf7835bdc3447c7a4fecec55)

3 weeks agomds: add MDS code to allow snap metadata mutations
Rishabh Dave [Mon, 17 Nov 2025 12:21:48 +0000 (17:51 +0530)]
mds: add MDS code to allow snap metadata mutations

Fixes: https://tracker.ceph.com/issues/66293
Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit a7617f354eb897ce86ba0b45eafa3005c3924d47)

3 weeks agomds: rename a finisher method in Server.cc since...
Rishabh Dave [Mon, 2 Mar 2026 16:16:32 +0000 (21:46 +0530)]
mds: rename a finisher method in Server.cc since...

it will be used in a different context as well in the upcoming commit.

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 047f41aac142bd70d07e25affed7d99cbd96320c)

3 weeks agomds: enable logging for snap.cc
Rishabh Dave [Sat, 14 Mar 2026 10:37:00 +0000 (16:07 +0530)]
mds: enable logging for snap.cc

Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit 2aed10be29fa7736051f03c66493d99606fe61e1)

3 weeks agoclient: add client code to allow snap metadata mutations
Rishabh Dave [Thu, 27 Nov 2025 14:10:27 +0000 (19:40 +0530)]
client: add client code to allow snap metadata mutations

Fixes: https://tracker.ceph.com/issues/66293
Signed-off-by: Rishabh Dave <ridave@redhat.com>
(cherry picked from commit bf5080c08cd86fe39976ebd5310506d7059dbfc6)

3 weeks agoqa: Add mgr snapshot mirror status tests
Kotresh HR [Tue, 23 Jun 2026 15:13:18 +0000 (20:43 +0530)]
qa: Add mgr snapshot mirror status tests

Add teuthology coverage for `ceph fs snapshot mirror status`:
parity with asok peer_status, default idle metrics, stale omap
handling, error paths, filter scopes, daemon restart, and cache TTL.

Reuse existing peer_dir_status and metrics assertion helpers.
Adjust stale test wait for InstanceWatcher timeout and set a
short cache TTL in the cache test. Pass peer_uuid via --peer_uuid=
in the test helper for peer-only scope queries.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit b80828fbadf285029239725d9cb175812aadf07a)

3 weeks agodoc/cephfs: Document fs snapshot mirror status mgr command
Kotresh HR [Sun, 21 Jun 2026 17:59:43 +0000 (23:29 +0530)]
doc/cephfs: Document fs snapshot mirror status mgr command

Describe the mirroring module command that reads persisted omap
metrics, including syntax, output layout, stale detection, caching,
and comparison with the admin socket peer status interface.

Document --peer_uuid as a named argument for peer-only and
directory/peer filtering.

Also update PendingReleaseNotes

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 87638c8be49918a8c794a1212f4122f2aa55f1c7)

3 weeks agomgr/mirroring: make snapshot mirror metrics cache optional
Kotresh HR [Sun, 21 Jun 2026 17:52:29 +0000 (23:22 +0530)]
mgr/mirroring: make snapshot mirror metrics cache optional

Add snapshot_mirror_metrics_cache_enabled (default true). When disabled,
metrics_status reads omap directly and skips complete and partial caches.
When enabled, behavior is unchanged.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 01dd9506f2c62a75841184ac9aa75cf72bf57b5b)

3 weeks agomgr/mirroring: make snapshot mirror metrics cache TTL configurable
Kotresh HR [Sun, 21 Jun 2026 17:50:28 +0000 (23:20 +0530)]
mgr/mirroring: make snapshot mirror metrics cache TTL configurable

Add snapshot_mirror_metrics_cache_ttl as a runtime mgr/mirroring module
option (default 15 seconds) instead of a hard-coded CACHE_TTL_SECS.
Both complete and partial lru_cache_timeout wrappers read the value
when caching omap metrics so operators can tune cache freshness without
code changes.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit e36faf8bcca5afd6ab4454751fb1d6e84a1bed5f)

3 weeks agomgr/mirroring: handle missing cephfs_mirror object in metrics status
Kotresh HR [Sun, 21 Jun 2026 17:48:15 +0000 (23:18 +0530)]
mgr/mirroring: handle missing cephfs_mirror object in metrics status

When snapshot mirroring is not enabled for a filesystem, the
cephfs_mirror RADOS object does not exist and "ceph fs snapshot mirror
status" fails reading sync stat omap. Map rados ENOENT to a clear
MirrorException so the CLI reports that snapshot mirroring must be
enabled instead of a generic omap read failure or an uncaught exception.

Also catch unexpected errors in metrics_status like other mirror CLI
handlers, so the mgr module does not crash on failure. Return
-errno.EINVAL from the generic error path instead of the exception
message as exit code.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit a7325bafb9ab29f17f45e4c14db56a86632d0ff1)

3 weeks agomgr/mirroring: Show default stats on newly added dir_root
Kotresh HR [Sun, 21 Jun 2026 17:35:30 +0000 (23:05 +0530)]
mgr/mirroring: Show default stats on newly added dir_root

When the directory is added for mirroring and the snapshot
is not taken yet, the peer_status show the following default
metrics.

{
  'state': 'idle',
  'snaps_synced': 0,
  'snaps_deleted': 0,
  'snaps_renamed': 0,
}

The mgr interface should also match that.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit c24b8d3cc2fb36ea77276598a2d1c2f90b1304ec)

3 weeks agomgr/mirroring: detect stale snapshot mirror sync metrics in omap
Kotresh HR [Sun, 21 Jun 2026 17:28:38 +0000 (22:58 +0530)]
mgr/mirroring: detect stale snapshot mirror sync metrics in omap

Persisted metrics in the cephfs_mirror omap can outlive the writing
daemon when cephfs-mirror stops or a directory is reshuffled to another
instance; mgr would keep reporting stale progress until the owning
daemon writes again.

Extend format_and_order_sync_stat_for_display() to compare persisted
_instance_id against InstanceWatcher live instances (via
FSPolicy.get_live_instance_ids()) and the directory's tracked instance
(via Policy.get_tracked_instance_id()). Mark metrics stale when the
persisted writer is no longer live (any state), or when it does not
match the tracked instance while persisted state is not "idle". Show
state "stale" with current_syncing_snap omitted.

Pass policy and live instance ids through load_sync_stat_metrics() and
fetch_sync_stat_metrics(), and into sync_stat_complete_cache and
sync_stat_partial_cache loaders. Cache hits serve already-formatted
(stale-marked) entries until TTL expiry without re-checking instance
liveness.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 15e998b865b99b3091054321cf0fdf6df9f327ab)

3 weeks agomgr/mirroring: cache fs snapshot mirror status omap metrics
Kotresh HR [Tue, 16 Jun 2026 10:42:41 +0000 (16:12 +0530)]
mgr/mirroring: cache fs snapshot mirror status omap metrics

Add a short-lived in-memory cache for metrics returned by
"ceph fs snapshot mirror status", which reads persisted sync stats
from the cephfs_mirror object omap. Omap walks are relatively
expensive; caching reduces repeated reads when the CLI or multiple
clients poll at short intervals.

Two TTL-bucketed LRU caches (CACHE_TTL_SECS) are used instead of one
unified cache. The complete cache always holds every mirrored
directory and peer for a filesystem; the partial cache holds one
directory. A single directory-granularity cache cannot prove it has
all directories, so full-scan queries rely on the complete cache
contract.

Cache implementation (lru_cache_timeout decorator backed by
_TimedLRUCache, not functools.lru_cache):

- lru_cache has no TTL and no peek-on-hit API. Single-directory
  queries must read the complete cache without loading on miss;
  lru_cache.cache is also unavailable on Python 3.14+.

- complete (sync_stat_complete_cache): full omap prefix scan via
  load_sync_stat_metrics. Cache key: (time_token, filesystem).
  COMPLETE_CACHE_MAX limits filesystem entries per TTL window.

- partial (sync_stat_partial_cache): per-directory omap key load via
  fetch_sync_stat_metrics. Cache key: (time_token, filesystem,
  dir_path, peer_ids). PARTIAL_CACHE_MAX limits directory entries.

Bind cache_peek/cache_info/cache_clear via a CachedMethod descriptor
so TTL lookups receive (self, ...); otherwise single-dir status fails
with 'str' object has no attribute 'mgr'.

Serve logic (under the existing lock):

- status <fs> [--peer_uuid=<uuid>]: load complete cache on miss;
  filter by peer when requested.

- status <fs> <dir>: peek complete cache (no load on miss); if the
  entry contains the directory and peers, serve from complete.
  Otherwise load partial cache (omap on miss).

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 8647c2ee79649b002c23bfb88757955640f886ce)

3 weeks agomgr/mirroring: Reorder and format metrics output
Kotresh HR [Mon, 1 Jun 2026 09:29:31 +0000 (14:59 +0530)]
mgr/mirroring: Reorder and format metrics output

Reorder and format status output to match the output of asok
interface peer_status command. Return metrics under
metrics/<dir>/peer/<uuid> to match the asok peer_status layout,
including {"metrics": {}} when no peers are configured.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 4c19167719caeab7de059bb1ba983e3f16afb5b3)

3 weeks agomgr/mirroring: Add new interface to expose mirroring metrics
Kotresh HR [Wed, 29 Apr 2026 09:19:15 +0000 (14:49 +0530)]
mgr/mirroring: Add new interface to expose mirroring metrics

Add the following new interface to expose mirroring metrics

ceph fs snapshot mirror status <fsname> [<mirrored_dir_path>] [--peer_uuid=<peer_uuid>]

The cmd loads the persisted directory sync metrics from the
cephfs_mirror object's omap. Metrics are grouped by mirrored
directory and peer.

When --peer_uuid is specified, only metrics for that peer are
returned. peer_uuid is a named CLI argument (_end_positional_) so
peer-only filtering does not require a mirrored directory path.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 33c572242ba78924df3e65f9fec23f2aefe81c66)

3 weeks agotools/cephfs_mirror: Remove persisted dir stats
Kotresh HR [Sat, 20 Jun 2026 07:58:56 +0000 (13:28 +0530)]
tools/cephfs_mirror: Remove persisted dir stats

When a directory is removed from mirroring, the persisted directory
stats need to be removed.  This patch handles the cleanup.

Omap keys must not be removed when mirrored directories are reshuffled
across cephfs-mirror daemons.  The mgr release notify now carries a
purging flag (set only during permanent removal, not reshuffle), and
the daemon removes persisted stats only when purging is true.  On
reshuffle with an in-progress sync, clear live current_syncing_snap
state and persist idle metrics so the acquiring daemon does not inherit
stale syncing omap entries.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 51a8348dfd681dfd446c347a82e581f2ed52a4c6)

3 weeks agotools/cephfs_mirror: Load persisted mirror metrics from omap
Kotresh HR [Sat, 20 Jun 2026 06:21:42 +0000 (11:51 +0530)]
tools/cephfs_mirror: Load persisted mirror metrics from omap

Load last_synced_snap metadata from the cephfs_mirror object omap on
PeerReplayer initialization and when a mirrored directory is added.
Live current_syncing_snap metrics are not restored; they are rebuilt
when synchronization starts.

When the daemon restarts after a snapshot was synced on the remote but
metrics were not yet written to omap, loaded metadata may belong to an
older snapshot.  Add reconcile_last_synced_snap() to compare against
the remote snap map, clear stale last-sync fields, and update
last_synced_snap id/name in memory.

Treat snaps_synced, snaps_deleted, and snaps_renamed as per-session
counters.  Do not load them from omap; they start at zero for each
daemon session and are still reported via the admin socket.  Persist
omap metrics unconditionally after reconcile so the mgr picks up the
new instance id and cleared session counters on restart.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 4b907eaf08869ef472b391774c421a68361a1f0a)

3 weeks agotools/cephfs_mirror: Persist metrics to omap
Kotresh HR [Sat, 20 Jun 2026 05:33:49 +0000 (11:03 +0530)]
tools/cephfs_mirror: Persist metrics to omap

Persist snapshot mirroring metrics to the cephfs_mirror object omap
for the mgr/mirroring module to support status command.

The tick thread only keeps omap up to date for in-progress syncs on
registered directories. Persist explicitly when stats change so omap
is updated before a directory unregisters:

 - snap delete/rename propagation (inc_deleted_snap / inc_renamed_snap)
 - after each successful snap sync (following set_last_synced_stat)
 - after sync_snaps failure (following _inc_failed_count)
 - after sync_perms failure (following _inc_failed_count)

Without the explicit call sites, omap can keep stale live or idle state
after sync completes or fails because directories leave m_registered
before the next tick.

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit c5dd50967036529c819ecdad7bb86705a15f61c6)

3 weeks agotools/cephfs_mirror: Adds capability to persist metrics
Kotresh HR [Sun, 17 May 2026 21:01:58 +0000 (02:31 +0530)]
tools/cephfs_mirror: Adds capability to persist metrics

Adds the capability to persist mirroring metrics.
The metrics are persisted in the omap of the cephfs-mirror
object. Metrics are persisted asynchronously.

Each mirrored directory path stores the corresponding
metrics as the value of a unique omap key representing
the mirrored directory. The omap key is as below.

sync_stat/<fsname>/<peer_uuid>/<mirrored_dir_path>

Fixes: https://tracker.ceph.com/issues/76686
Signed-off-by: Kotresh HR <khiremat@redhat.com>
(cherry picked from commit 9f59332b35259773783bbd38e4f92fa1520d6cbf)