Backup Recovery   «Prev  Next»

Lesson 3Understanding Fast-Start checkpointing
ObjectiveDescribe how checkpoint parameters and commands affect Oracle instance recovery.

Understanding Fast-Start Checkpointing in Oracle 26ai

Checkpointing affects how much work Oracle must reconstruct after an instance fails. By writing modified database blocks to datafiles and advancing recovery progress, Oracle can reduce the outstanding work needed to return the database to service. The benefit has a cost: more demanding recovery targets can require additional writes during normal operation.

Oracle AI Database 26ai provides time-based checkpoint management through FAST_START_MTTR_TARGET. Administrators also need to understand redo-distance limits, time-based limits, explicit checkpoint commands, and the recovery estimates that show what the system is doing. These controls influence instance recovery; they do not replace an RMAN backup or restore lost storage.

The practical objective is to choose an achievable recovery target while maintaining acceptable application performance. A small configured number is useful only when the system can meet it. This lesson connects checkpoint progress to recovery, explains the documented controls, and shows how to observe their effects without confusing an estimate with a measured outage.

How Database Changes Reach Persistent Storage

Oracle processes database blocks in the buffer cache[2]. A block modified in memory becomes a dirty buffer[1] because its current contents differ from the corresponding block in the datafile. Database writer processes, identified as DBWn and often described collectively as DBWR, write these modified blocks to datafiles.

Changes also generate redo records. Log writer, LGWR, writes redo from the redo log buffer to online redo logs. A commit depends on the appropriate redo becoming durable, rather than on every changed data block being written immediately. Conversely, DBWn can write a modified block whose transaction has not yet committed. Physical block persistence and transaction commitment are different events.

The checkpoint process, CKPT, coordinates checkpoint information and updates the relevant file metadata. It does not perform LGWR's redo-writing role. Checkpointing should therefore not be described as copying transactions into redo logs, waiting for a full redo log buffer, or clearing that buffer for the next batch of transactions.

In ARCHIVELOG mode, archiver processes copy completed online redo logs to archive destinations. This supports recovery procedures that require archived redo. A log switch, redo writing, archiving, and checkpoint progress are related activities, but they have separate purposes. Keeping these responsibilities distinct makes both recovery planning and performance diagnosis easier.

Checkpoint Position and Recovery Work

After an instance failure, Oracle performs cache recovery, also called the roll-forward phase[3], by applying required redo to affected blocks. The reconstructed changes can include committed and uncommitted work. Transaction recovery then uses undo to remove changes belonging to unfinished transactions.

Checkpoint position identifies recovery progress. As older modified buffers reach the datafiles, the position can advance and less earlier work remains to be reconstructed. Avoid defining it merely as the last committed transaction or assuming that every change after it is uncommitted. Recovery follows physical change records and transaction state, which are not interchangeable.

Consider two instances of the same workload observed at different times. In the first observation, substantial changes remain only in memory, backed by durable redo. In the second, background writes have persisted more of those changes. If a failure occurs, the second situation can require less reconstruction. The actual recovery duration still depends on storage, redo processing, startup overhead, and the work present at failure.

Incremental checkpointing advances progress during ordinary activity instead of relying entirely on large bursts of work. This reduces the amount of outstanding recovery work that can accumulate. More frequent or aggressive writing can, however, compete with application I/O. The correct configuration balances expected recovery effort against the resources consumed while the database is healthy.

Checkpointing cannot reconstruct a missing datafile by itself. Media recovery may require restoring a backup and applying redo to that restored content. An instance crash with intact files and a storage failure that destroys files therefore require different responses. Determine which failure occurred before choosing a recovery procedure.

Checkpoint Parameters in Oracle 26ai

Controls, units, and recovery effects
ControlMeaningOperational consideration
FAST_START_MTTR_TARGETTarget in seconds for single-instance crash recovery.Use a practical positive target; configured and effective targets can differ.
LOG_CHECKPOINT_INTERVALLimits redo distance using physical operating-system blocks.A meaningful nonzero value overrides target-based management; zero ignores this limit.
LOG_CHECKPOINT_TIMEOUTConstrains incremental checkpoint age in seconds.Its documented default is 1800; zero disables time-based checkpoints.
LOG_CHECKPOINTS_TO_ALERTControls checkpoint messages in the alert log.Provides diagnostic information rather than a recovery-time target.
ALTER SYSTEM CHECKPOINTRequests an explicit checkpoint.Waits for completion and can add write activity.

Set an Achievable MTTR Target

FAST_START_MTTR_TARGET has a documented range of 0 through 3600 seconds and a default of 0. It is dynamically modifiable with ALTER SYSTEM, is not modifiable in a PDB, and can have different values on different RAC instances. A positive setting supports recovery-target management; zero should not be interpreted as a promise of immediate recovery.

Oracle's performance tuning guide describes the target as the expected time to start the instance and perform cache recovery. It therefore includes database startup work rather than cache recovery alone. It does not measure an operating-system reboot, restoring damaged files, completing every background rollback, reconnecting clients, or validating the application.

Choose a value inside the practical range for the actual environment. Startup overhead creates a lower bound; workload and the amount of potentially outstanding recovery work affect the upper bound. Setting a target below what the system can achieve can increase checkpointing without producing the requested startup time. A less demanding value may reduce normal write pressure, but its benefit also depends on other active limits.

The following is an illustrative configuration command for a test environment, not a universal production recommendation:

ALTER SYSTEM SET FAST_START_MTTR_TARGET = 30;

The administrator needs the appropriate privilege and container context. Before using a command in a real deployment, determine the existing value, which instance should change, and whether the change should persist. The effect of an omitted SCOPE depends on whether the instance uses an SPFILE or a text initialization file. Choose the intended scope explicitly in the operational procedure rather than assuming the example is temporary.

Lower targets can shorten expected recovery by requiring more writes before a failure. Higher targets can give ordinary activity more breathing room. Neither direction is automatically best. A service with strict availability needs may accept additional normal I/O; a workload already limited by storage may suffer if checkpoint activity increases without an accompanying capacity improvement.

Understand the Other Limits Before Combining Settings

LOG_CHECKPOINT_INTERVAL concerns the redo distance between the incremental checkpoint and the latest redo written. Its unit is physical operating-system blocks, not database blocks or minutes. Its default of zero removes this parameter's limit. A nonzero setting can override FAST_START_MTTR_TARGET; log switches also cause checkpoint activity.

LOG_CHECKPOINT_TIMEOUT concerns elapsed checkpoint age and the time buffers remain dirty. The default is 1800 seconds. Setting it to zero disables time-based checkpoints, which Oracle discourages unless an MTTR target is set. It is misleading to describe this as a simple timer that always forces an independent full checkpoint at fixed intervals.

When using the MTTR target, Oracle's tuning guide instructs administrators to disable or remove competing LOG_CHECKPOINT_INTERVAL, LOG_CHECKPOINT_TIMEOUT, and historical FAST_START_IO_TARGET settings. Review existing configuration before changing these controls. Adding a time target while leaving stricter legacy limits active can make the actual writing behavior differ from the intended policy.

FAST_START_IO_TARGET belongs to the historical I/O-based approach. The current recovery view retains an obsolete compatibility column for it that is always null. Do not copy the legacy 1000-buffer example into a new configuration or assume an old parameter remains usable. Verify supported settings for the installed release and document any migration from historical configuration.

The names FAST_START_MTTR, CHECKPOINT_PROCESS, and CHECKPOINT_TIMEOUT in older lesson text are not the documented checkpoint controls described here. Use the exact supported names. A configuration review should distinguish a valid parameter, an obsolete setting, and a typographical error before considering performance implications.

Observe Recovery Estimates and Checkpoint Writes

V$INSTANCE_RECOVERY exposes recovery estimates and information about checkpoint mechanisms. These illustrative read-only queries require access to the relevant dynamic performance view. Query in the appropriate administrative context, and record the instance and observation time:

SELECT target_mttr, estimated_mttr,
       recovery_estimated_ios, actual_redo_blks,
       writes_mttr, writes_autotune
FROM v$instance_recovery;

TARGET_MTTR is the effective target; ESTIMATED_MTTR reflects the estimated recovery duration for current work. Dirty-buffer and redo-block information provides additional context. Write counters indicate activity attributed to target-based management and automatic checkpoint tuning. Their values describe mechanisms, not the elapsed duration of a recovery that has actually been performed.

Take observations across representative workload periods. An update burst can increase the estimate while background writes catch up. One sample above the target does not establish persistent failure to meet the recovery objective. Conversely, one quiet sample below it does not demonstrate acceptable behavior during the busiest part of the day.

For cumulative counters, compare changes over a known interval rather than treating a total as a current rate. Record elapsed time, update volume, and relevant storage conditions. This allows a meaningful comparison between configurations and helps distinguish ordinary workload changes from a change in checkpoint policy.

RAC diagnosis must also retain instance identity. A local view describes the instance being queried; an appropriate global view can support comparisons across the cluster. Different configured targets or uneven workload placement may produce different observations. Do not average away an instance whose service workload is the actual subject of the availability requirement.

Force a Checkpoint Deliberately

ALTER SYSTEM CHECKPOINT requests a checkpoint and does not return until it completes. Oracle requires an open database and the ALTER SYSTEM privilege. The operation makes committed changes persistent in datafiles, but should not be confused with committing every session's pending transaction.

ALTER SYSTEM CHECKPOINT;

In RAC, the documented default is GLOBAL, which checkpoints the instances that have opened the database. LOCAL limits the operation to the redo thread associated with the issuing instance. These variants make the intended scope explicit:

ALTER SYSTEM CHECKPOINT LOCAL;
ALTER SYSTEM CHECKPOINT GLOBAL;

These are alternatives illustrating scope, not a sequence to execute routinely. Repeated forced checkpoints can consume resources needed by applications. Completion also does not freeze the database: further updates can create additional dirty buffers[4] and new recovery work immediately afterward.

A checkpoint is neither an independent backup nor a substitute for archive management. It does not establish that client failover works or that a business transaction completed successfully. Use it where an operational procedure requires it, and interpret its outcome within that procedure's actual purpose.

Evaluate the Tradeoff with Representative Work

A useful comparison starts with a concrete question: can the service meet its availability requirement during a busy update period without an unacceptable increase in normal response time? A database observed only during quiet hours provides little evidence about that question. Include the transaction patterns and data access that matter to the application, and keep the assessment's scope consistent between trials.

For example, an order-processing system may perform frequent small commits alongside an overnight batch. The batch can change the volume and age of modified blocks substantially. A checkpoint policy assessed only against daytime traffic may produce a different recovery estimate during the batch. Record both conditions and decide which outage scenario the target is intended to address. The business requirement should guide the choice of representative workload.

Diagnostic logging can help correlate checkpoints with other activity. If checkpoint messages are enabled, use their timestamps together with workload and storage observations. Do not infer that enabling more messages makes checkpoint progress faster. Likewise, a large number of checkpoint messages does not independently establish a storage bottleneck: examine the work occurring at the same time before attributing a performance change to checkpointing.

Keep a small assessment record containing the original settings, proposed settings, observation interval, application measurements, recovery measurements, and the reason for the final choice. This makes the decision reviewable by another administrator. It also provides a baseline for later changes rather than leaving a target value whose rationale has been forgotten.

Start a checkpoint-tuning assessment by defining the milestone that matters. Instance open time is one milestone; first successful service request and completion of affected transactions are others. Agree on the required behavior before choosing a target. Otherwise, a setting may appear successful while users still experience an unacceptable interruption.

Record the baseline configuration, database release, normal workload, and storage characteristics. Observe recovery estimates and ordinary I/O under load. Select one practical target change, assess the corresponding behavior, and retain a way to restore the previous configuration. Changing several interacting limits at once makes it harder to explain an improvement or regression.

Deliberate crash and recovery calibration belongs in a designated test environment under an approved procedure. Preserve diagnostic records and confirm the expected committed outcomes after recovery. Compare similar workloads and failure conditions; cached storage after a process failure can behave differently from storage after a wider hardware outage.

Checkpoint tuning is successful when it supports the required recovery behavior with acceptable normal performance. Keep measured results separate from estimates and vendor feature descriptions. Revisit the assessment when workload, memory, storage, or availability requirements materially change.

Check Your Understanding

Before continuing, explain why a more recent checkpoint can reduce reconstruction work, why a lower MTTR target can increase normal writes, and why an explicit checkpoint cannot replace a backup. Identify which metric is an estimate and which service milestone requires a real recovery test.

Fast-Start Fault Recovery Quiz

Continue with Fast-Start on-demand rollback to examine how Oracle handles requests for blocks affected by unfinished transactions after cache recovery.

[1] Dirty buffer: A database block modified in memory whose current contents have not yet been written to its datafile.
[2] Buffer cache: Shared memory used to hold database blocks for access and modification.
[3] Roll-forward phase: Cache recovery applies required redo to affected blocks, including committed and uncommitted changes, before transaction recovery removes unfinished changes.
[4] Dirty buffers after a checkpoint: Continuing database activity can modify blocks again, creating new persistence and recovery work.

Oracle 26ai References


SEMrush Software 3 SEMrush Banner 3