2026-09-20 13:14:54 Displayed 47 times

ext4fs: a plan for journalling support

https://git.kmx.io/IABSD.fr/src/_tree/master/sys/ufs/ext4fs/plan.md

Goal

Add reliable JBD2 journalling support to IABSD's ext4fs implementation. The first supported runtime mode will be metadata-only data=ordered journalling with a single filesystem-wide transaction. More advanced modes and concurrency can be added after recovery and crash consistency are proven.

Current state

The tree already contains mount-time JBD2 replay in sys/ufs/ext4fs/ext4fs_journal.c. It implements a basic three-pass scan, revoke, and replay flow. Runtime filesystem operations do not write JBD2 transactions, however; metadata buffers still reach their home locations directly through bwrite(), bdwrite(), and bawrite().

The existing replay implementation also needs hardening before it is safe to use as the recovery side of a journal writer. In particular, it assumes a mostly linear descriptor -> data -> commit transaction and does not validate all modern JBD2 checksums and features.

Phase 1: Harden journal recovery

Phase 1 tests

Phase 2: Introduce the runtime journal core

Start with a serialized implementation: one running transaction and one committing transaction per mounted filesystem. Correctness is more important than batching or throughput in this phase.

Add journal state to struct m_ext4fs, including:

Introduce an interface along these lines:

int  ext4fs_journal_begin(struct mount *, unsigned int,
    struct ext4fs_journal_handle **);
int  ext4fs_journal_get_write_access(struct ext4fs_journal_handle *,
    struct buf *, u_int64_t);
int  ext4fs_journal_dirty_metadata(struct ext4fs_journal_handle *,
    struct buf *);
int  ext4fs_journal_revoke(struct ext4fs_journal_handle *, u_int64_t);
int  ext4fs_journal_end(struct ext4fs_journal_handle *);
int  ext4fs_journal_force_commit(struct mount *);
void ext4fs_journal_abort(struct mount *, int);

Each operation must reserve enough journal credits before modifying metadata. Initial credit estimates can be conservative while the implementation is serialized.

Phase 3: Implement ordered commits

Implement metadata-only data=ordered commits in this order:

flush affected regular-file data to home locations
    -> write descriptor blocks, metadata payloads, and revoke blocks
    -> issue a durability flush
    -> write the commit block
    -> issue a durability flush
    -> checkpoint metadata to its home locations
    -> issue a durability flush
    -> advance the journal tail and persist journal state

Additional requirements:

Phase 4: Convert metadata writers

Audit every metadata write in sys/ufs/ext4fs and route it through a journal handle. This includes:

Direct bwrite(), bdwrite(), or bawrite() calls must remain only for regular-file data, the journal's own I/O, recovery, checkpointing, or another explicitly documented exception.

Wrap each compound namespace operation in one transaction:

Phase 5: VFS semantics

Phase 6: Crash-consistency test matrix

Use filesystem images created by Linux tools and run IABSD in a VM. Inject an abrupt power loss after each commit phase and at journal wraparound boundaries.

For every recovered image:

  1. boot or remount it on IABSD;
  2. verify expected namespace and file-data outcomes;
  3. run e2fsck -fn;
  4. mount it on Linux and repeat integrity checks;
  5. confirm replay is idempotent by attempting recovery again.

Exercise at least:

Deferred work

Do not include these in the first working milestone:

References

Update (2026-09-30)

The plan has been completed. We get 1.24Mb/s throughput with modern NVMe SSDs on ext4fs with journaling on IABSD.

After a first step of optimization and debug we still hit a 30Mb/s limit bulk throughput and some deadlocks in getblk.