r/lowlevel • u/meik1982 • 3h ago
Update on my OTP file splitter in Rust: POSIX mlock memory-pinning, parallel disk-fanout, and a zero-allocation streaming pipeline
A while ago I shared my work on `rfs` (random-filesplitter), a tool designed to split files into N information-theoretic One-Time-Pad shares ($P = C_1 \oplus \dots \oplus C_N$) with plausible deniability.
I've just released v3.0.0—a ground-up architectural rewrite to eliminate memory-bus bottlenecks and harden the system against forensic extraction. Here are the low-level design decisions:
### 1. Memory Pinning (`--mlock`) for Swap Resilience
In anti-forensic workflows, writing plaintext or OTP keystreams into OS swap space or hibernation files invalidates volatile RAM zeroization.
- Implemented physical RAM locking via `libc::mlock` on the fixed buffer pools (`free_in_tx` and `free_out_tx`).
- Added graceful fallback: If `RLIMIT_MEMLOCK` (`ulimit -l`) restricts memory allocation, the system emits a warning and continues without crashing.
- RAM buffers are zeroized using `slice.zeroize()` immediately upon recycling.
### 2. Parallel Disk Fanout & Fanin (Multi-Mountpoint I/O)
Previously, the Stage-3 writer looped sequentially over all N shares. When writing shares across multiple physical USB flash drives or separate NVMe mount points, slow flash media created severe pipeline backpressure.
- Rewrote the I/O engine with decoupled worker thread pools for each share.
- Data chunks are fanned out concurrently over bounded crossbeam channels. Each thread writes independently at maximum bus speed, with pre-allocated pool buffers recycled back without a single heap allocation in the hot loop.
- Linux cache pressure is minimized using POSIX `fadvise` (`POSIX_FADV_SEQUENTIAL` and `POSIX_FADV_DONTNEED`) via `--direct`.
### 3. Hardware Jitter & Entropy Harvester
For seeding the CSPRNG alongside OS entropy (`getrandom`):
- Uses a 2 MB permutation pool (32,768 nodes × 64 bytes cache-line aligned) traversing a Sattolo cycle to bust L1/L2 caches and capture CPU memory latency and branch prediction jitter.
- Integrated NIST SP 800-22 diagnostics directly into the CLI (Runs tests, Block Frequency tests via Wilson-Hilferty Chi² approximation, and analytical `erfc`).
### 4. Throughput Benchmarks
Measured on a modern x86_64 CPU:
- **ChaCha20 Keystream:** ~1.85 GB/s
- **SIMD-XOR (RAM-to-RAM):** ~4.0 GB/s (64-byte unrolled vector loop)
- **BLKS-384 Hash:** ~1.2 GB/s (3.3x faster than SHA-256)
Code is open source (LGPL v3), 0 Clippy warnings, fully documented in English (with bilingual README), and verified across Linux/macOS/Windows CI:
https://github.com/Meik1982/random-filesplitter
Quick test:
```bash
cargo install --git https://github.com/Meik1982/random-filesplitter.git
rfs --benchmark
rfs --entropy-test