| Match messages in logs (every line would be required to be present in log output Copy from "Messages before crash" column below): | |
| Match messages in full crash (every line would be required to be present in crash log output Copy from "Full Crash" column below): | |
| Limit to a test: (Copy from below "Failing text"): | |
| Delete these reports as invalid (real bug in review or some such) | |
| Bug or comment: | |
| Extra info: |
| Failing Test | Full Crash | Messages before crash | Link and date |
|---|---|---|---|
| replay-dual test 26: dbench and tar with mds failover | LustreError: 150553:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 150553:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 1 PID: 150553 Comm: umount Kdump: loaded Tainted: G O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x2e6/0x3b0 [obdclass] class_cleanup+0x410/0xac0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f69f229f8fb | Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 1 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffff8aa24593d000: operation ldlm_enqueue to node 0@lo failed: rc = -19 LustreError: Skipped 1 previous similar message Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1233 to 0x2c0000400:1249) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1233 to 0x280000400:1249) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1234 to 0x240000400:1249) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1233 to 0x300000400:1249) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 2 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 4 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1320 to 0x240000400:1345) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1320 to 0x300000400:1345) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1321 to 0x280000400:1345) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1320 to 0x2c0000400:1345) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc LustreError: 140116:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. LustreError: 140116:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 8 previous similar messages Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1434 to 0x300000400:1473) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1434 to 0x2c0000400:1473) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1434 to 0x280000400:1473) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1434 to 0x240000400:1473) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 4 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 9 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:02, of 2 clients 2 recovered and 0 were evicted. Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1495 to 0x240000400:1537) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1495 to 0x2c0000400:1537) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1496 to 0x300000400:1537) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1495 to 0x280000400:1537) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 5 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) LustreError: 140756:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 0000000077766c15 ns: mdt-lustre-MDT0000_UUID lock: ffff8aa22ae38600/0x571985fad36946e1 lrc: 3/0,0 mode: CW/CW res: [0x20000afe1:0x3d:0x0].0x0 bits 0x5/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x571985fad36946a9 expref: 3 pid: 140756 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 LustreError: 117999:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff8aa26cd72680 x1875436216001024/t0(0) o6->lustre-OST0002-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-2-0.0' uid:0 gid:0 projid:4294967295 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 14 previous similar messages Lustre: lustre-MDT0000 is waiting for obd_unlinked_exports more than 8 seconds. The obd refcount = 3. Is it stuck? Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc LustreError: 117985:0:(client.c:3449:ptlrpc_replay_interpret()) @@@ status 301, old was 0 req@ffff8aa229f69f80 x1875436211085312/t94489280696(94489280696) o101->lustre-MDT0000-mdc-ffff8aa25b1da000@0@lo:12/10 lens 648/608 e 0 to 0 dl 1788556782 ref 2 fl Interpret:RPQU/604/0 rc 301/301 job:'dbench.0' uid:0 gid:0 projid:0 LustreError: 117985:0:(client.c:3449:ptlrpc_replay_interpret()) Skipped 470 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1555 to 0x240000400:1601) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1555 to 0x280000400:1601) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1555 to 0x300000400:1601) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1556 to 0x2c0000400:1601) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 6 times Lustre: Failing over lustre-MDT0000 LustreError: 141539:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 0000000069898d02 ns: mdt-lustre-MDT0000_UUID lock: ffff8aa25aa6c000/0x571985fad36a2d4e lrc: 3/0,0 mode: PR/PR res: [0x20000afe1:0x1b4:0x0].0x0 bits 0x1b/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x571985fad36a2d39 expref: 3 pid: 141539 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 3 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 50 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1663 to 0x2c0000400:1697) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1664 to 0x240000400:1697) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1664 to 0x300000400:1697) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1663 to 0x280000400:1697) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 7 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 11 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000-lwp-OST0003: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete Lustre: Skipped 60 previous similar messages Lustre: 117999:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1788556817/real 1788556817] req@ffff8aa26cd7b100 x1875436218017280/t0(0) o400->lustre-MDT0000-lwp-OST0000@0@lo:12/10 lens 224/224 e 0 to 1 dl 1788556833 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 Lustre: 117999:0:(client.c:2504:ptlrpc_expire_one_request()) Skipped 74 previous similar messages Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1770 to 0x300000400:1793) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1771 to 0x2c0000400:1793) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1770 to 0x280000400:1793) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1770 to 0x240000400:1793) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 8 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc LustreError: MGC192.168.123.73@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail LustreError: Skipped 10 previous similar messages Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect Lustre: Skipped 11 previous similar messages Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect Lustre: Skipped 11 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1864 to 0x2c0000400:1889) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1865 to 0x240000400:1889) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1865 to 0x280000400:1889) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1865 to 0x300000400:1889) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 9 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc LustreError: 143866:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. LustreError: 143866:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 3 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1901 to 0x240000400:1921) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1901 to 0x280000400:1921) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1902 to 0x300000400:1921) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1901 to 0x2c0000400:1921) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 10 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffff8aa24593d000: operation mds_reint to node 0@lo failed: rc = -19 LustreError: Skipped 17 previous similar messages Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 2 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1940 to 0x2c0000400:1985) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1940 to 0x300000400:1985) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1940 to 0x280000400:1985) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1939 to 0x240000400:1985) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 11 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2045 to 0x240000400:2081) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2045 to 0x300000400:2081) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2045 to 0x280000400:2081) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2045 to 0x2c0000400:2081) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 12 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2130 to 0x2c0000400:2145) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2130 to 0x240000400:2177) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2130 to 0x280000400:2145) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2130 to 0x300000400:2177) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 13 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2214 to 0x280000400:2241) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2246 to 0x240000400:2273) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2214 to 0x2c0000400:2241) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2247 to 0x300000400:2273) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 14 times Lustre: Failing over lustre-MDT0000 LustreError: 117990:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff8aa25acf1500 x1875436224545280/t0(0) o6->lustre-OST0001-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-1-0.0' uid:0 gid:0 projid:4294967295 LustreError: 117990:0:(client.c:1394:ptlrpc_import_delay_req()) Skipped 1 previous similar message Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2310 to 0x240000400:2337) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2310 to 0x300000400:2337) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2278 to 0x2c0000400:2305) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2278 to 0x280000400:2305) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 15 times Lustre: Failing over lustre-MDT0000 LustreError: 147038:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000bd05cda4 ns: mdt-lustre-MDT0000_UUID lock: ffff8aa26b214400/0x571985fad3745a04 lrc: 3/0,0 mode: PR/PR res: [0x20000afe1:0x2e0:0x0].0x0 bits 0x1b/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x571985fad37459da expref: 3 pid: 147038 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 6 previous similar messages LustreError: 117998:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff8aa1d0683800 x1875436225358208/t0(0) o6->lustre-OST0003-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-3-0.0' uid:0 gid:0 projid:4294967295 LustreError: 117998:0:(client.c:1394:ptlrpc_import_delay_req()) Skipped 2 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2359 to 0x240000400:2401) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2327 to 0x280000400:2369) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2360 to 0x300000400:2401) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2327 to 0x2c0000400:2369) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 16 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2400 to 0x280000400:2433) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2432 to 0x300000400:2465) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2432 to 0x240000400:2465) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2400 to 0x2c0000400:2433) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 17 times Lustre: Failing over lustre-MDT0000 LustreError: 148381:0:(ldlm_lib.c:1199:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. LustreError: 148381:0:(ldlm_lib.c:1199:target_handle_connect()) Skipped 3 previous similar messages Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2534 to 0x240000400:2561) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2534 to 0x300000400:2561) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2502 to 0x280000400:2529) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2502 to 0x2c0000400:2529) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 18 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2628 to 0x300000400:2657) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2595 to 0x2c0000400:2625) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2628 to 0x240000400:2657) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2596 to 0x280000400:2625) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 19 times Lustre: Failing over lustre-MDT0000 LustreError: 149840:0:(ldlm_resource.c:1207:ldlm_resource_complain()) mdt-lustre-MDT0000_UUID: namespace resource [0x2000013a1:0xf6a:0x0].0x0 (ffff8aa258fccc00) refcount nonzero (1) after lock cleanup; forcing cleanup. Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:2680 to 0x2c0000400:2721) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:2681 to 0x280000400:2721) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:2712 to 0x240000400:2753) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:2712 to 0x300000400:2753) Lustre: DEBUG MARKER: centos-71.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 20 times Lustre: Failing over lustre-MDT0000 LustreError: 140550:0:(ldlm_lockd.c:2564:ldlm_cancel_handler()) ldlm_cancel from 0@lo arrived at 1788557178 with bad export cookie 6276194868053733458 LustreError: 117990:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff8aa2789cad80 x1875436230295424/t0(0) o6->lustre-OST0003-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-3-0.0' uid:0 gid:0 projid:4294967295 LustreError: 117990:0:(client.c:1394:ptlrpc_import_delay_req()) Skipped 17 previous similar messages LustreError: 150038:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000b3cfd021 ns: mdt-lustre-MDT0000_UUID lock: ffff8aa1ce72d400/0x571985fad37a4ca7 lrc: 4/0,0 mode: CW/CW res: [0x20000afe1:0x54d:0x0].0x0 bits 0x5/0x0 rrc: 3 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x571985fad37a4c84 expref: 4 pid: 150038 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-09-04 21:26) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 343441:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 343441:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 10 PID: 343441 Comm: umount Kdump: loaded Tainted: G O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] ? string+0x60/0x80 cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xac0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7fb8eb3208fb | Lustre: DEBUG MARKER: Started rundbench load pid=342671 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 15 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3963 to 0x240000400:4001) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3918 to 0x2c0000400:3937) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3918 to 0x280000400:3937) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3919 to 0x300000400:3937) Lustre: DEBUG MARKER: centos-31.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) LustreError: 343053:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000000c275c8d ns: mdt-lustre-MDT0000_UUID lock: ffff8a6c9d55be00/0xcca06d024d43994a lrc: 4/0,0 mode: PR/PR res: [0x20001a9e3:0xf92:0x0].0x0 bits 0x1b/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0xcca06d024d43993c expref: 3 pid: 343053 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-31 23:07) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 81899:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 81899:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 13 PID: 81899 Comm: umount Kdump: loaded Tainted: G O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xac0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7efd38a698fb | Lustre: DEBUG MARKER: Started rundbench load pid=75501 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 6 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 16 previous similar messages Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect Lustre: Skipped 19 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3961 to 0x240000400:4001) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3917 to 0x2c0000400:3937) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3948 to 0x280000400:3969) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3917 to 0x300000400:3937) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3991 to 0x280000400:4033) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3960 to 0x2c0000400:4001) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3955 to 0x300000400:4001) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4017 to 0x240000400:4033) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000-mdc-ffff92b40f98f000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete Lustre: Skipped 33 previous similar messages Lustre: server umount lustre-MDT0000 complete LustreError: MGC192.168.123.53@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail LustreError: Skipped 6 previous similar messages Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 34 previous similar messages Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4052 to 0x280000400:4097) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4019 to 0x2c0000400:4065) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4022 to 0x300000400:4065) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4043 to 0x240000400:4065) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec LustreError: 77969:0:(osd_handler.c:720:osd_ro()) lustre-MDT0000: *** setting device osd-zfs read-only *** LustreError: 77969:0:(osd_handler.c:720:osd_ro()) Skipped 4 previous similar messages Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 4 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffff92b40f98f000: operation ldlm_enqueue to node 0@lo failed: rc = -19 LustreError: Skipped 5 previous similar messages LustreError: 3285:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff92b46cacad80 x1875007542182528/t0(0) o6->lustre-OST0001-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-1-0.0' uid:0 gid:0 projid:4294967295 LustreError: 3285:0:(client.c:1394:ptlrpc_import_delay_req()) Skipped 4 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects Lustre: Skipped 9 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4090 to 0x240000400:4129) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4089 to 0x2c0000400:4129) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4122 to 0x280000400:4161) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4084 to 0x300000400:4129) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 5 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4143 to 0x2c0000400:4161) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4145 to 0x300000400:4161) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4176 to 0x280000400:4193) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4140 to 0x240000400:4161) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 6 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 1 previous similar message Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4187 to 0x300000400:4225) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4211 to 0x280000400:4257) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4186 to 0x2c0000400:4225) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4183 to 0x240000400:4225) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 7 times Lustre: Failing over lustre-MDT0000 LustreError: 80048:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000d243834a ns: mdt-lustre-MDT0000_UUID lock: ffff92b469b44800/0x97923223b6695d6a lrc: 3/0,0 mode: CW/CW res: [0x20001a9e3:0xee5:0x0].0x0 bits 0x5/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x97923223b6695d5c expref: 4 pid: 80048 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: 3287:0:(client.c:2504:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1788149267/real 1788149267] req@ffff92b4625b3800 x1875007543195520/t0(0) o400->lustre-MDT0000-lwp-OST0002@0@lo:12/10 lens 224/224 e 0 to 1 dl 1788149283 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 Lustre: 3287:0:(client.c:2504:ptlrpc_expire_one_request()) Skipped 53 previous similar messages Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 4 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4273 to 0x280000400:4289) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4246 to 0x300000400:4289) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4244 to 0x2c0000400:4289) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4235 to 0x240000400:4257) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 8 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4309 to 0x280000400:4353) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4309 to 0x300000400:4353) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4310 to 0x2c0000400:4353) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4280 to 0x240000400:4321) Lustre: DEBUG MARKER: centos-51.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 9 times Lustre: Failing over lustre-MDT0000 LustreError: 81472:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000003993d3a7 ns: mdt-lustre-MDT0000_UUID lock: ffff92b456351a00/0x97923223b66ab209 lrc: 4/0,0 mode: CW/CW res: [0x20001a9e3:0x118f:0x0].0x0 bits 0x5/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x97923223b66ab1fb expref: 4 pid: 81472 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-31 04:09) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 2500630:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 2500630:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 12 PID: 2500630 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] ? string+0x60/0x80 cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f6e8db1e8fb | Lustre: DEBUG MARKER: Started rundbench load pid=2498683 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 16 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3962 to 0x240000400:4001) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3917 to 0x280000400:3937) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3918 to 0x300000400:3937) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3918 to 0x2c0000400:3937) Lustre: DEBUG MARKER: centos-36.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 LustreError: 2499173:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000009b73f03f ns: mdt-lustre-MDT0000_UUID lock: ffffa0f0cec90000/0x92d6b8ecb471119 lrc: 3/0,0 mode: CW/CW res: [0x20001a9e3:0xedd:0x0].0x0 bits 0x5/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x92d6b8ecb47110b expref: 4 pid: 2499173 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3963 to 0x280000400:4001) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3963 to 0x2c0000400:4001) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3963 to 0x300000400:4001) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4027 to 0x240000400:4065) Lustre: DEBUG MARKER: centos-36.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 4 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 33 previous similar messages Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4020 to 0x280000400:4065) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4084 to 0x240000400:4129) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4019 to 0x2c0000400:4065) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4020 to 0x300000400:4065) Lustre: DEBUG MARKER: centos-36.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 4 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffffa0f0b4140800: operation ldlm_cancel to node 0@lo failed: rc = -19 LustreError: Skipped 6 previous similar messages LustreError: 2500244:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000d0a970ec ns: mdt-lustre-MDT0000_UUID lock: ffffa0f0e1142c00/0x92d6b8ecb489c4c lrc: 4/0,0 mode: CW/CW res: [0x20001a9e3:0x1034:0x0].0x0 bits 0x5/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x92d6b8ecb489c3e expref: 4 pid: 2500244 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-30 02:56) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 3565920:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 3565630:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 0000000083b7f334 ns: mdt-lustre-MDT0000_UUID lock: ffff932dc524dc00/0x48013b3dc4ca58f4 lrc: 4/0,0 mode: CW/CW res: [0x20001a9e3:0x1207:0x0].0x0 bits 0x5/0x0 rrc: 3 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x48013b3dc4ca58e6 expref: 4 pid: 3565630 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 LustreError: 3565920:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 12 PID: 3565920 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] ? do_raw_spin_unlock+0x79/0x190 cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f16965068fb | Lustre: DEBUG MARKER: Started rundbench load pid=3558760 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 LustreError: 3486088:0:(client.c:1394:ptlrpc_import_delay_req()) @@@ IMP_CLOSED req@ffff932d83269180 x1874890953587072/t0(0) o6->lustre-OST0002-osc-MDT0000@0@lo:28/4 lens 544/432 e 0 to 0 dl 0 ref 1 fl Rpc:QU/200/ffffffff rc 0/-1 job:'osp-syn-2-0.0' uid:0 gid:0 projid:4294967295 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 4 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 16 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3959 to 0x240000400:4033) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3915 to 0x280000400:3937) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3915 to 0x2c0000400:3937) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3915 to 0x300000400:3937) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4053 to 0x240000400:4097) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3957 to 0x300000400:4001) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3956 to 0x2c0000400:4001) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3956 to 0x280000400:4001) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000-lwp-OST0001: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete Lustre: Skipped 33 previous similar messages Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 34 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4020 to 0x280000400:4065) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4117 to 0x240000400:4161) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4020 to 0x300000400:4065) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4020 to 0x2c0000400:4065) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec LustreError: 3561206:0:(osd_handler.c:720:osd_ro()) lustre-MDT0000: *** setting device osd-zfs read-only *** LustreError: 3561206:0:(osd_handler.c:720:osd_ro()) Skipped 4 previous similar messages Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 4 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4185 to 0x240000400:4225) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4089 to 0x2c0000400:4129) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4090 to 0x280000400:4129) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4089 to 0x300000400:4129) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 5 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffff932dbd532800: operation ldlm_enqueue to node 0@lo failed: rc = -19 LustreError: Skipped 3 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects Lustre: Skipped 9 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4242 to 0x240000400:4257) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4145 to 0x300000400:4161) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4145 to 0x2c0000400:4161) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4145 to 0x280000400:4161) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 6 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: 3563205:0:(ldlm_lib.c:2976:target_recovery_thread()) too long recovery - read logs LustreError: dumping log to /tmp/lustre-log.1788038137.3563205 Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4184 to 0x300000400:4225) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4184 to 0x280000400:4225) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4184 to 0x2c0000400:4225) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4280 to 0x240000400:4321) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 7 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete LustreError: MGC192.168.123.93@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail LustreError: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4346 to 0x240000400:4385) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4250 to 0x2c0000400:4289) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4251 to 0x280000400:4289) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4250 to 0x300000400:4289) Lustre: 3486091:0:(client.c:2503:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1788038144/real 1788038144] req@ffff932d7ae5f380 x1874890956821760/t0(0) o400->lustre-MDT0000-lwp-OST0003@0@lo:12/10 lens 224/224 e 0 to 1 dl 1788038160 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 Lustre: 3486091:0:(client.c:2503:ptlrpc_expire_one_request()) Skipped 70 previous similar messages Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 8 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LustreError: 3486085:0:(client.c:3448:ptlrpc_replay_interpret()) @@@ status 301, old was 0 req@ffff932db8f0f380 x1874890953362432/t317827580346(317827580346) o101->lustre-MDT0000-mdc-ffff932dbd532800@0@lo:12/10 lens 576/608 e 0 to 0 dl 1788038201 ref 2 fl Interpret:RPQU/604/0 rc 301/301 job:'dbench.0' uid:0 gid:0 projid:0 LustreError: 3486085:0:(client.c:3448:ptlrpc_replay_interpret()) Skipped 326 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4304 to 0x280000400:4321) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4401 to 0x240000400:4417) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4305 to 0x2c0000400:4321) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4305 to 0x300000400:4321) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 9 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4345 to 0x2c0000400:4385) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4345 to 0x280000400:4385) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4441 to 0x240000400:4481) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4345 to 0x300000400:4385) Lustre: DEBUG MARKER: centos-91.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 10 times Lustre: Failing over lustre-MDT0000 | Link to test (2026-08-29 21:17) |
| replay-single test 70b: dbench 3mdts recovery; 1 clients | LustreError: 1662531:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 1662531:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 1 PID: 1662531 Comm: umount Kdump: loaded Tainted: G O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f046f9818fb | Lustre: DEBUG MARKER: Started rundbench load pid=1661209 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete LDISKFS-fs (dm-0): 6 truncates cleaned up LDISKFS-fs (dm-0): recovery complete LDISKFS-fs (dm-0): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc LustreError: MGC192.168.123.33@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail LustreError: Skipped 6 previous similar messages Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 3 clients reconnect Lustre: Skipped 9 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:05, of 3 clients 3 recovered and 0 were evicted. Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000402:3748 to 0x2c0000402:3777) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000402:3706 to 0x300000402:3745) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x380000402:3706 to 0x380000402:3745) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x340000402:3706 to 0x340000402:3745) Lustre: DEBUG MARKER: centos-31.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds2 REPLAY BARRIER on lustre-MDT0001 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0001 Lustre: DEBUG MARKER: test_70b fail mds2 2 times Lustre: Failing over lustre-MDT0001 Lustre: server umount lustre-MDT0001 complete LustreError: lustre-MDT0001-mdc-ffff8b6318a4f000: operation mds_statfs to node 0@lo failed: rc = -107 LustreError: Skipped 5 previous similar messages LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) lustre-MDT0001-mdc-ffff8b6318a4f000: can't stat MDS #0: rc = -107 LustreError: 1605152:0:(ldlm_lib.c:1190:target_handle_connect()) lustre-MDT0001: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. LustreError: 1605152:0:(ldlm_lib.c:1190:target_handle_connect()) Skipped 119 previous similar messages LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) lustre-MDT0001-mdc-ffff8b6318a4f000: can't stat MDS #0: rc = -19 LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) lustre-MDT0001-mdc-ffff8b6318a4f000: can't stat MDS #0: rc = -19 LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) lustre-MDT0001-mdc-ffff8b6318a4f000: can't stat MDS #0: rc = -19 LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) lustre-MDT0001-mdc-ffff8b6318a4f000: can't stat MDS #0: rc = -19 LustreError: 1661234:0:(lmv_obd.c:1471:lmv_statfs()) Skipped 1 previous similar message LDISKFS-fs (dm-1): 12 truncates cleaned up LDISKFS-fs (dm-1): recovery complete LDISKFS-fs (dm-1): mounted filesystem with ordered data mode. Opts: user_xattr,errors=remount-ro,no_mbcache,nodelalloc Lustre: lustre-MDT0001: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 6 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x2c0000401:204 to 0x2c0000401:225) Lustre: lustre-OST0001: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x300000401:202 to 0x300000401:225) Lustre: lustre-OST0003: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x380000401:203 to 0x380000401:225) Lustre: lustre-OST0002: new connection from lustre-MDT0001-mdtlov (cleaning up unused objects from 0x340000401:203 to 0x340000401:225) Lustre: DEBUG MARKER: centos-31.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0001-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0001-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds3 REPLAY BARRIER on lustre-MDT0002 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0002 Lustre: DEBUG MARKER: test_70b fail mds3 3 times Lustre: Failing over lustre-MDT0002 LustreError: 1606211:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 0000000028bb2444 ns: mdt-lustre-MDT0002_UUID lock: ffff8b630263b800/0x5fb42379a3449023 lrc: 4/0,0 mode: PR/PR res: [0x2800007ed:0x3bb:0x0].0x0 bits 0x1b/0x0 rrc: 3 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x5fb42379a3449015 expref: 3 pid: 1606211 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-29 09:47) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 481260:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 481260:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 4 PID: 481260 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] cleanup_resource+0x30d/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7fb74a4d98fb | Lustre: DEBUG MARKER: Started rundbench load pid=480279 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 LustreError: 477917:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000002d8d751a ns: mdt-lustre-MDT0000_UUID lock: ffff95d21ba85c00/0x24ce7d8cc33bcdeb lrc: 5/0,0 mode: PR/PR res: [0x20001a9e3:0xf1c:0x0].0x0 bits 0x1b/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x24ce7d8cc33bcddd expref: 3 pid: 477917 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 16 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3991 to 0x240000400:4033) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3947 to 0x2c0000400:3969) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3947 to 0x280000400:3969) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3946 to 0x300000400:3969) Lustre: DEBUG MARKER: centos-86.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 LustreError: 480833:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000007c940f68 ns: mdt-lustre-MDT0000_UUID lock: ffff95d21a168200/0x24ce7d8cc33c59b3 lrc: 4/0,0 mode: PR/PR res: [0x20001a9e3:0xf59:0x0].0x0 bits 0x1b/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x24ce7d8cc33c59a5 expref: 3 pid: 480833 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-28 12:51) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 313960:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 313536:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 000000009197c908 ns: mdt-lustre-MDT0000_UUID lock: ffff9eaa3a796000/0x57943875b988f945 lrc: 4/0,0 mode: PR/PR res: [0x20001a9e3:0x10b6:0x0].0x0 bits 0x1b/0x0 rrc: 3 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x57943875b988f937 expref: 3 pid: 313536 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 LustreError: 313960:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 8 PID: 313960 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] ? string+0x60/0x80 cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f74cf1f28fb | Lustre: DEBUG MARKER: Started rundbench load pid=309922 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000-mdc-ffff9eaaacebb000: Connection to lustre-MDT0000 (at 0@lo) was lost; in progress operations using this service will wait for recovery to complete Lustre: Skipped 30 previous similar messages LustreError: 307851:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 0000000080e9742c ns: mdt-lustre-MDT0000_UUID lock: ffff9eaa0a59a200/0x57943875b985aba3 lrc: 3/0,0 mode: PR/PR res: [0x20001a9e3:0xf1d:0x0].0x0 bits 0x1b/0x0 rrc: 2 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x57943875b985ab95 expref: 3 pid: 307851 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: server umount lustre-MDT0000 complete LustreError: MGC192.168.123.98@tcp: Connection to MGS (at 0@lo) was lost; in progress operations using this service will fail LustreError: Skipped 5 previous similar messages Lustre: lustre-MDT0000: Imperative Recovery not enabled, recovery window 60-180 Lustre: Skipped 19 previous similar messages Lustre: lustre-MDT0000: in recovery but waiting for the first client to connect Lustre: Skipped 8 previous similar messages Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 32 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:3959 to 0x240000400:4001) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3915 to 0x2c0000400:3937) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3914 to 0x280000400:3937) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3915 to 0x300000400:3937) Lustre: DEBUG MARKER: centos-96.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 2 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3954 to 0x280000400:3969) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4020 to 0x240000400:4065) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:3959 to 0x2c0000400:4001) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:3956 to 0x300000400:4001) Lustre: 237238:0:(client.c:2503:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1787701251/real 1787701251] req@ffff9eaa8cb81880 x1874537772435328/t0(0) o400->lustre-MDT0000-lwp-OST0000@0@lo:12/10 lens 224/224 e 0 to 1 dl 1787701267 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 Lustre: 237238:0:(client.c:2503:ptlrpc_expire_one_request()) Skipped 40 previous similar messages Lustre: DEBUG MARKER: centos-96.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec LustreError: 311595:0:(osd_handler.c:715:osd_ro()) lustre-MDT0000: *** setting device osd-zfs read-only *** LustreError: 311595:0:(osd_handler.c:715:osd_ro()) Skipped 4 previous similar messages Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects Lustre: Skipped 9 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:01, of 1 clients 1 recovered and 0 were evicted. Lustre: Skipped 9 previous similar messages Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4086 to 0x240000400:4129) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:3988 to 0x280000400:4033) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4019 to 0x2c0000400:4065) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4026 to 0x300000400:4065) Lustre: DEBUG MARKER: centos-96.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 4 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4155 to 0x240000400:4193) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4058 to 0x280000400:4097) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4091 to 0x300000400:4129) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4091 to 0x2c0000400:4129) Lustre: DEBUG MARKER: centos-96.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 5 times Lustre: Failing over lustre-MDT0000 LustreError: lustre-MDT0000-mdc-ffff9eaaacebb000: operation ldlm_enqueue to node 0@lo failed: rc = -19 LustreError: Skipped 7 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:4146 to 0x2c0000400:4161) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:4153 to 0x300000400:4193) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:4126 to 0x280000400:4161) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:4215 to 0x240000400:4257) Lustre: DEBUG MARKER: centos-96.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 6 times Lustre: Failing over lustre-MDT0000 | Link to test (2026-08-25 23:42) |
| replay-dual test 26: dbench and tar with mds failover | LustreError: 3064482:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 3064482:0:(ldlm_lockd.c:1018:ldlm_server_completion_ast()) LBUG CPU: 0 PID: 3064482 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x2e6/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f9345f9e8fb | Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 1 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: server umount lustre-MDT0000 complete LustreError: 3034485:0:(client.c:3448:ptlrpc_replay_interpret()) @@@ status 301, old was 0 req@ffff99e018ed7700 x1874479966450688/t94489280716(94489280716) o101->lustre-MDT0000-mdc-ffff99e024cae800@0@lo:12/10 lens 648/608 e 0 to 0 dl 1787644808 ref 2 fl Interpret:RPQU/604/0 rc 301/301 job:'dbench.0' uid:0 gid:0 projid:0 LustreError: 3034485:0:(client.c:3448:ptlrpc_replay_interpret()) Skipped 2 previous similar messages Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1349 to 0x280000400:1377) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1318 to 0x240000400:1345) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1222 to 0x2c0000400:1249) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:998 to 0x300000400:1025) Lustre: DEBUG MARKER: centos-106.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 2 times Lustre: Failing over lustre-MDT0000 LustreError: 3061632:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000f97417bf ns: mdt-lustre-MDT0000_UUID lock: ffff99e06e802200/0x54845280a64aa413 lrc: 3/0,0 mode: CW/CW res: [0x20000afe1:0x3a:0x0].0x0 bits 0x5/0x0 rrc: 2 type: IBT gid 0 flags: 0x50200000000000 nid: 0@lo remote: 0x54845280a64aa405 expref: 4 pid: 3061632 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 1 previous similar message Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 4 previous similar messages Lustre: server umount lustre-MDT0000 complete Lustre: lustre-MDT0000-lwp-OST0001: Connection restored to 0@lo (at 0@lo) Lustre: Skipped 32 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1301 to 0x2c0000400:1345) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1403 to 0x240000400:1441) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1078 to 0x300000400:1121) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1421 to 0x280000400:1441) Lustre: DEBUG MARKER: centos-106.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 3 times Lustre: Failing over lustre-MDT0000 Lustre: server umount lustre-MDT0000 complete Lustre: 3034489:0:(client.c:2503:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1787644853/real 1787644853] req@ffff99e01693d080 x1874479968313984/t0(0) o400->lustre-MDT0000-lwp-OST0000@0@lo:12/10 lens 224/224 e 0 to 1 dl 1787644869 ref 1 fl Rpc:XNQr/200/ffffffff rc 0/-1 job:'kworker.0' uid:0 gid:0 projid:4294967295 Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 2 clients reconnect Lustre: 3034489:0:(client.c:2503:ptlrpc_expire_one_request()) Skipped 57 previous similar messages Lustre: Skipped 7 previous similar messages Lustre: lustre-MDT0000: Recovery over after 0:02, of 2 clients 2 recovered and 0 were evicted. Lustre: Skipped 8 previous similar messages Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1395 to 0x2c0000400:1441) Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1503 to 0x240000400:1537) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1489 to 0x280000400:1537) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1165 to 0x300000400:1185) Lustre: DEBUG MARKER: centos-106.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 4 times Lustre: Failing over lustre-MDT0000 Lustre: lustre-MDT0000: Not available for connect from 0@lo (stopping) Lustre: Skipped 5 previous similar messages LustreError: 3063168:0:(ldlm_lib.c:1190:target_handle_connect()) lustre-MDT0000: not available for connect from 0@lo (no target). If you are running an HA pair check that the target is mounted on the other server. Lustre: server umount lustre-MDT0000 complete Lustre: lustre-OST0000: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x240000400:1597 to 0x240000400:1633) Lustre: lustre-OST0001: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x280000400:1588 to 0x280000400:1633) Lustre: lustre-OST0002: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x2c0000400:1496 to 0x2c0000400:1537) Lustre: lustre-OST0003: new connection from lustre-MDT0000-mdtlov (cleaning up unused objects from 0x300000400:1250 to 0x300000400:1281) Lustre: DEBUG MARKER: centos-106.localnet: executing wait_import_state_mount (FULL|IDLE) mdc.lustre-MDT0000-mdc-*.mds_server_uuid 1475 0 Lustre: DEBUG MARKER: mdc.lustre-MDT0000-mdc-*.mds_server_uuid in FULL state after 0 sec Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_26 fail mds1 5 times Lustre: Failing over lustre-MDT0000 LustreError: 3064109:0:(ldlm_lockd.c:1453:ldlm_handle_enqueue()) ### lock on destroyed export 00000000e2940198 ns: mdt-lustre-MDT0000_UUID lock: ffff99e007941c00/0x54845280a64d65a7 lrc: 4/0,0 mode: CW/CW res: [0x20000afe1:0x11c:0x0].0x0 bits 0x5/0x0 rrc: 3 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x54845280a64d6599 expref: 4 pid: 3064109 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-25 08:01) |
| replay-single test 70b: dbench 1mdts recovery; 1 clients | LustreError: 109802:0:(ldlm_lockd.c:1058:ldlm_server_completion_ast()) ASSERTION( data != ((void *)0) ) failed: LustreError: 109802:0:(ldlm_lockd.c:1058:ldlm_server_completion_ast()) LBUG CPU: 9 PID: 109802 Comm: umount Kdump: loaded Tainted: G W O -------- - - 4.18.0rocky8.10-debug #2 Hardware name: Red Hat KVM, BIOS 1.16.0-4.module+el8.9.0+1408+7b966129 04/01/2014 Call Trace: dump_stack+0x99/0xca lbug_with_loc.cold.4+0xd/0x86 [libcfs] ldlm_server_completion_ast+0x635/0xd80 [ptlrpc] ? string+0x60/0x80 cleanup_resource+0x30d/0x490 [ptlrpc] ldlm_resource_clean+0x3f/0x70 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_relax+0x267/0x5f0 [obdclass] ? cleanup_resource+0x490/0x490 [ptlrpc] ? cleanup_resource+0x490/0x490 [ptlrpc] cfs_hash_for_each_nolock+0x1a0/0x2c0 [obdclass] ldlm_namespace_cleanup+0x38/0xf0 [ptlrpc] __ldlm_namespace_free+0x72/0x6a0 [ptlrpc] ? _cond_resched+0x21/0x50 ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 ldlm_namespace_free_prior+0x87/0x2c0 [ptlrpc] mdt_device_fini+0x233/0xb70 [mdt] obd_precleanup.isra.17+0xb3/0x380 [obdclass] ? class_disconnect_exports+0x1d1/0x3b0 [obdclass] class_cleanup+0x410/0xab0 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? _raw_spin_unlock+0x16/0x30 class_process_config+0xdb0/0x2450 [obdclass] class_manual_cleanup+0x4b0/0xa60 [obdclass] ? do_raw_spin_unlock+0x79/0x190 ? class_name2obd+0x142/0x190 [obdclass] server_put_super+0x11dc/0x1cb0 [ptlrpc] ? fsnotify_grab_connector+0x5b/0xb0 ? fsnotify_sb_delete+0x234/0x320 generic_shutdown_super+0xb7/0x1c0 kill_anon_super+0x20/0x50 lustre_kill_super+0x2e/0x60 [lustre] deactivate_locked_super+0x59/0xd0 deactivate_super+0x88/0xa0 cleanup_mnt+0x63/0xe0 __cleanup_mnt+0x1a/0x30 task_work_run+0xd2/0x120 exit_to_usermode_loop+0x1e4/0x200 do_syscall_64+0x3de/0x3f0 entry_SYSCALL_64_after_hwframe+0x49/0xae RIP: 0033:0x7f6997fa28fb | Lustre: DEBUG MARKER: Started rundbench load pid=109598 ... Lustre: DEBUG MARKER: mds1 REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: local REPLAY BARRIER on lustre-MDT0000 Lustre: DEBUG MARKER: test_70b fail mds1 1 times Lustre: Failing over lustre-MDT0000 LustreError: 107259:0:(ldlm_lockd.c:1509:ldlm_handle_enqueue()) ### lock on destroyed export 00000000c22ccbde ns: mdt-lustre-MDT0000_UUID lock: ffff9e1e2941b400/0x85192323b794a97d lrc: 4/0,0 mode: CW/CW res: [0x20001a9e3:0xf24:0x0].0x0 bits 0x5/0x0 rrc: 4 type: IBT gid 0 flags: 0x50306400000000 nid: 0@lo remote: 0x85192323b794a96f expref: 4 pid: 107259 timeout: 0 lvb_type: 0 lru_score: 0 lru_type: 0 | Link to test (2026-08-22 14:22) |