联系:手机/微信(+86 17813235971) QQ(107644445)
标题:kcratr_nab_less_than_odr和system坏块故障处理
作者:惜分飞©版权所有[未经本人同意,不得以任何形式转载,否则有进一步追究法律责任的权利.]
学校客户由于机房断电,导致数据库异常.现场工程师进行了现场备份,然后尝试强制拉库,结果没有成功,大概报错和操作过程如下:
断电之后启动数据库,数据库报ORA-600 kcratr_nab_less_than_odr故障解决错误,这个是一种非常常见的操作,一般是由于写丢失导致,以前有过很多类似恢复经历:
ORA-600 kcratr_nab_less_than_odr故障解决
差点被误操作的ORA-600 kcratr_nab_less_than_odr故障
ORA-600 kcratr_nab_less_than_odr和ORA-600 2662故障处理
ORA-600 kcratr_nab_less_than_odr和ORA-600 4193故障处理
Tue Aug 18 23:59:19 2026 ALTER DATABASE OPEN This instance was first to open Beginning crash recovery of 2 threads parallel recovery started with 19 processes Started redo scan Completed redo scan read 21634 KB redo, 800 data blocks need recovery Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_275950.trc (incident=703037): ORA-00600: internal error code, arguments: [kcratr_nab_less_than_odr], [1], [6916], [3], [4], [], [], [], [] Incident details in: /u01/app/oracle/diag/rdbms/orcl/orcl2/incident/incdir_703037/orcl2_ora_275950_i703037.trc Use ADRCI or Support Workbench to package the incident. See Note 411.1 at My Oracle Support for error and packaging details. Abort recovery for domain 0 Aborting crash recovery due to error 600 Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_275950.trc: ORA-00600: internal error code, arguments: [kcratr_nab_less_than_odr], [1], [6916], [3], [4], [], [], [], [] Abort recovery for domain 0 Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_275950.trc: ORA-00600: internal error code, arguments: [kcratr_nab_less_than_odr], [1], [6916], [3], [4], [], [], [], [] ORA-600 signalled during: ALTER DATABASE OPEN...
现场恢复人员上来之后,直接尝试做强制resetlogs操作,数据库报ORA-600 ORA-600 krsi_al_hdr_update.15错误,主要是由于redo写丢失导致无法resetlogs成功,具体参考:Alter Database Open Resetlogs returns error ORA-00600: [krsi_al_hdr_update.15], (Doc ID 2026541.1) Oracle断电故障处理
Wed Aug 19 00:28:47 2026 alter database open resetlogs RESETLOGS is being done without consistancy checks. This may result in a corrupted database. The database should be recreated. RESETLOGS after incomplete recovery UNTIL CHANGE 29721342127 Archived Log entry 38935 added for thread 1 sequence 6916 ID 0xceea62af dest 1: ARCH: All Archive destinations made inactive due to error 742 ARCH: Closing local archive destination LOG_ARCHIVE_DEST_1: '+ARCHDG/2_19951_1224781365.arc' (error 742)(orcl2) Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_287355.trc (incident=727038): ORA-00600: internal error code, arguments: [krsi_al_hdr_update.15], [4294967295], [], [], [], [], [], [], [] Incident details in: /u01/app/oracle/diag/rdbms/orcl/orcl2/incident/incdir_727038/orcl2_ora_287355_i727038.trc Use ADRCI or Support Workbench to package the incident. See Note 411.1 at My Oracle Support for error and packaging details. Master archival failure: 600 Archive all online redo logfiles failed:600 ORA-600 signalled during: alter database open resetlogs...
通过 ALTER DATABASE RECOVER database using backup controlfile until cancel之后,继续尝试强制打开库,报ORA-600 2662错误.
Wed Aug 19 00:40:25 2026 Checker run found 32 new persistent data failures alter database open resetlogs RESETLOGS is being done without consistancy checks. This may result in a corrupted database. The database should be recreated. RESETLOGS after incomplete recovery UNTIL CHANGE 29721342127 Archived Log entry 38936 added for thread 1 sequence 6915 ID 0xceea62af dest 1: Archived Log entry 38937 added for thread 1 sequence 6916 ID 0xceea62af dest 1: Archived Log entry 38938 added for thread 2 sequence 19951 ID 0xceea62af dest 1: Archived Log entry 38939 added for thread 2 sequence 19950 ID 0xceea62af dest 1: Clearing online redo logfile 1 +DATADG/orcl/onlinelog/group_1.319.1224781365 Clearing online log 1 of thread 1 sequence number 6915 Wed Aug 19 00:40:34 2026 Clearing online redo logfile 1 complete Clearing online redo logfile 2 +DATADG/orcl/onlinelog/group_2.320.1224781367 Clearing online log 2 of thread 1 sequence number 6916 Clearing online redo logfile 2 complete Clearing online redo logfile 3 +DATADG/orcl/onlinelog/group_3.323.1224781463 Clearing online log 3 of thread 2 sequence number 19951 Clearing online redo logfile 3 complete Clearing online redo logfile 4 +DATADG/orcl/onlinelog/group_4.324.1224781465 Clearing online log 4 of thread 2 sequence number 19950 Clearing online redo logfile 4 complete Resetting resetlogs activation ID 3471467183 (0xceea62af) Online log +DATADG/orcl/onlinelog/group_1.319.1224781365: Thread 1 Group 1 was previously cleared Online log +ARCHDG/orcl/onlinelog/group_1.4574.1224781367: Thread 1 Group 1 was previously cleared Online log +DATADG/orcl/onlinelog/group_2.320.1224781367: Thread 1 Group 2 was previously cleared Online log +ARCHDG/orcl/onlinelog/group_2.10729.1224781369: Thread 1 Group 2 was previously cleared Online log +DATADG/orcl/onlinelog/group_3.323.1224781463: Thread 2 Group 3 was previously cleared Online log +ARCHDG/orcl/onlinelog/group_3.12866.1224781463: Thread 2 Group 3 was previously cleared Online log +DATADG/orcl/onlinelog/group_4.324.1224781465: Thread 2 Group 4 was previously cleared Online log +ARCHDG/orcl/onlinelog/group_4.7527.1224781465: Thread 2 Group 4 was previously cleared Wed Aug 19 00:40:43 2026 Setting recovery target incarnation to 3 Wed Aug 19 00:40:43 2026 Assigning activation ID 3488382552 (0xcfec7e58) Thread 2 opened at log sequence 1 Current log# 3 seq# 1 mem# 0: +DATADG/orcl/onlinelog/group_3.323.1224781463 Current log# 3 seq# 1 mem# 1: +ARCHDG/orcl/onlinelog/group_3.12866.1224781463 Successful open of redo thread 2 MTTR advisory is disabled because FAST_START_MTTR_TARGET is not set Wed Aug 19 00:40:43 2026 SMON: enabling cache recovery Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_291405.trc (incident=739007): ORA-00600: internal error code, arguments: [2662], [6], [3951555809], [6], [3951556555], [12583040], [], [] Incident details in: /u01/app/oracle/diag/rdbms/orcl/orcl2/incident/incdir_739007/orcl2_ora_291405_i739007.trc Wed Aug 19 00:40:45 2026 Use ADRCI or Support Workbench to package the incident. See Note 411.1 at My Oracle Support for error and packaging details. Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_291405.trc: ORA-00600: internal error code, arguments: [2662], [6], [3951555809], [6], [3951556555], [12583040], [], [] Errors in file /u01/app/oracle/diag/rdbms/orcl/orcl2/trace/orcl2_ora_291405.trc: ORA-00600: internal error code, arguments: [2662], [6], [3951555809], [6], [3951556555], [12583040], [], [] Error 600 happened during db open, shutting down database USER (ospid: 291405): terminating the instance due to error 600 Instance terminated by USER, pid = 291405 ORA-1092 signalled during: alter database open resetlogs...
到这一步,现场停止了继续尝试,我接手故障处理.先dbv检测坏块,发现system有两个坏块
[oracle@db3 ~]$ dbv userid=sys/oracle file=/datapool/cold_backup_20260818/SYSTEM.314.1224781265 DBVERIFY: Release 11.2.0.4.0 - Production on Sat Aug 22 10:39:26 2026 Copyright (c) 1982, 2011, Oracle and/or its affiliates. All rights reserved. DBVERIFY - Verification starting : FILE = /datapool/cold_backup_20260818/SYSTEM.314.1224781265 Page 94587 is marked corrupt Corrupt block relative dba: 0x0041717b (file 1, block 94587) Bad header found during dbv: Data in bad block: type: 11 format: 2 rdba: 0x00400001 last change scn: 0x0000.00000000 seq: 0x1 flg: 0x04 spare1: 0x0 spare2: 0x0 spare3: 0x0 consistency value in tail: 0x00000b01 check value in block header: 0xd49 computed block checksum: 0x0 Page 95021 is marked corrupt Corrupt block relative dba: 0x0041732d (file 1, block 95021) Bad header found during dbv: Data in bad block: type: 11 format: 2 rdba: 0x00400001 last change scn: 0x0000.00000000 seq: 0x1 flg: 0x04 spare1: 0x0 spare2: 0x0 spare3: 0x0 consistency value in tail: 0x00000b01 check value in block header: 0xd49 computed block checksum: 0x0 DBVERIFY - Verification complete Total Pages Examined : 157440 Total Pages Processed (Data) : 72079 Total Pages Failing (Data) : 0 Total Pages Processed (Index): 21854 Total Pages Failing (Index): 0 Total Pages Processed (Other): 49484 Total Pages Processed (Seg) : 1 Total Pages Failing (Seg) : 0 Total Pages Empty : 14021 Total Pages Marked Corrupt : 2 Total Pages Influx : 0 Total Pages Encrypted : 0 Highest block SCN : 0 (0.0)
OBET> dbv file 1 =============================================== DBV (Data Block Verification) Block Size: 8192 bytes Endian: little-endian (x86) Target File: #1(only) =============================================== Verifying file #1: /datapool/cold_backup_20260818/SYSTEM.314.1224781265 (157441 blocks) - Started: 2026-08-22 11:14:48 File #32: rfile=1 (0x00000001) header_block_num=157440 (0x00026700) filesize_status:OK Progress: 100000 / 157441 blocks checked... File #1completed: 0 all zero, 0 soft corrupted, 2 tailchk error, 0 checksum error, 0 rdba error DBV completed at: 2026-08-22 11:14:59 =============================================== DBV Summary: Total blocks checked: 157439 Total all zero blocks found: 0 Total all rdba error blocks found: 0 Total all tailchk error blocks found: 2 Total all soft corrupted blocks found: 0 Total all checksum error blocks found: 0 Total bad blocks found: 2 Execution time: 11.00 seconds Throughput: 111.82 MB/s =============================================== Detailed report saved to: dbv_file_32_20260822111448.log Files processed: 1 OBET> list corrupt file 1(/datapool/cold_backup_20260818/SYSTEM.314.1224781265) total bad blocks: 2 block# bad block type 94587 tailchk 95021 tailchk
使用obet修复坏块
Oracle Block Editor Tool使用手册
OBET> set file 32 filename set to: /datapool/cold_backup_20260818/SYSTEM.314.1224781265 (file#1) OBET> set block 94587 block set to: 94587 OBET> set mode edit mode set to: edit OBET> repair block Warning: Missing value for 'block', using global setting: 94587 Repairing block 94587 in file /datapool/cold_backup_20260818/SYSTEM.314.1224781265... Repair analysis for block 94587: 1. seq_kcbh check: 0x01 -> OK 2. Tailchk check: 0x0106C0E4 -> needs repair (0xE4C00601) 3. Checksum check: 0x5EC5 -> needs repair (0x7DE6) Confirm repair operations: File: /datapool/cold_backup_20260818/SYSTEM.314.1224781265 Block: 94587 Operations needed: fix tailchk, fix checksum Confirm? (Y/YES to proceed): y Verification after repair: 1. seq_kcbh: 0x01 OK 2. Tailchk: 0x0106C0E4 OK 3. Checksum: 0x7DE6 OK Block 94587 repair completed successfully. OBET> set block 95021 block set to: 95021 OBET> repair block Warning: Missing value for 'block', using global setting: 95021 Repairing block 95021 in file /datapool/cold_backup_20260818/SYSTEM.314.1224781265... Repair analysis for block 95021: 1. seq_kcbh check: 0x01 -> OK 2. Tailchk check: 0xF8FA0601 -> needs repair (0xE06C0601) 3. Checksum check: 0x81BA -> OK Confirm repair operations: File: /datapool/cold_backup_20260818/SYSTEM.314.1224781265 Block: 95021 Operations needed: fix tailchk Confirm? (Y/YES to proceed): y Verification after repair: 1. seq_kcbh: 0x01 OK 2. Tailchk: 0x01066CE0 OK 3. Checksum: 0x81BA OK Block 95021 repair completed successfully.
dbv检查确认坏块修复成功
[oracle@db3 ~]$ dbv file=/datapool/cold_backup_20260818/SYSTEM.314.1224781265 DBVERIFY: Release 11.2.0.4.0 - Production on Sat Aug 22 11:17:30 2026 Copyright (c) 1982, 2011, Oracle and/or its affiliates. All rights reserved. DBVERIFY - Verification starting : FILE = /datapool/cold_backup_20260818/SYSTEM.314.1224781265 DBVERIFY - Verification complete Total Pages Examined : 157440 Total Pages Processed (Data) : 72080 Total Pages Failing (Data) : 0 Total Pages Processed (Index): 21855 Total Pages Failing (Index): 0 Total Pages Processed (Other): 49484 Total Pages Processed (Seg) : 1 Total Pages Failing (Seg) : 0 Total Pages Empty : 14021 Total Pages Marked Corrupt : 0 Total Pages Influx : 0 Total Pages Encrypted : 0 Highest block SCN : 3951556890 (6.3951556890)
后面的恢复比较简单,使用客户恢复之前的备份,直接重建ctl,然后打开库成功,并且做rman校验没有异常,直接把恢复之后的库备份还原到asm里面,完成本次恢复任务
