Correct scanner accounting and verify release packaging #2

Open
gamertan wants to merge 9 commits from sizequeen-scan-hardening into main
9 changed files with 839 additions and 179 deletions
Showing only changes of commit 8b177e21a0 - Show all commits
+18 -1
View File
@@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
### Fixed
- Count hard-link allocation once in a deterministic traversal order while
retaining apparent sizes per pathname. Filtered comparison has independent
allocation ownership, including when the first link is ignored.
- Report filesystem errors and incomplete ancestor totals instead of silently
dropping errors or treating vanished entries as accessible empty directories.
- Inherit ignore rules when scanning nested directories; accept dangling
symlinks as scan targets and distinguish special files from directories.
### Added
- A scanner module with structured reports, file identities, progress counters,
cooperative cancellation, and per-scan thread pools.
- Optional `-x` / `--one-file-system` boundary handling. Partial CLI reports
return a nonzero exit status and include `complete` in JSON output.
## [1.0.0] - 2026-01-22
This stable release marks the official launch of `sized` under the GNU General Public License v3.0.
@@ -19,4 +36,4 @@ This stable release marks the official launch of `sized` under the GNU General P
- **Exporting**: Support for CSV and NDJSON output formats.
- **Documentation**: Comprehensive standard Unix man pages and a detailed user guide.
- **Shell Completions**: Automatic generation for Bash, Zsh, and Fish.
- **Licensing**: GPL-3.0 to ensure the tool remains a shared public resource.
- **Licensing**: GPL-3.0 to ensure the tool remains a shared public resource.
+18 -3
View File
@@ -35,9 +35,24 @@ sized [path]
```
### Key Concepts
- **Disk Usage** (Default): The actual physical space consumed on disk (Blocks * 512 bytes). This is the "true" footprint and the metric `sized` prioritizes.
- **Disk Usage** (Default): Filesystem-reported allocation (Blocks * 512 bytes), with hard links counted once per scan. Shared extents and snapshots can make this differ from the space recovered by deletion.
- **Apparent Size**: The logical size of the entry (file length). Available via the `-a` or `--apparent` flag.
- **Blocks**: The actual filesystem blocks allocated.
- **Blocks**: Filesystem-reported 512-byte units of allocation. Directory metadata is included.
Symlinks are scanned as leaves, including dangling links; their targets are not
followed. Missing targets or unreadable entries are reported on stderr. Partial
scans return a nonzero exit status and JSON includes `complete: false`; inspect
that status before relying on a total.
### Filesystem Boundary (`-x` / `--one-file-system`)
Skip entries whose device differs from the scan root, useful when a directory
contains mounted filesystems. Boundary entries contribute no size, and JSON
marks them with `skipped_mount`.
```bash
sized -x /path/to/volume
```
### Path Display
- **Relative Path** (Default): `sized` shows paths relative to the current directory.
@@ -53,7 +68,7 @@ sized -d 1
```
> [!NOTE]
> Regardless of the display depth, `sized` always calculates the *total* size of all subdirectories accurately by traversing the entire tree.
> Display depth limits output, not scanning. The tree is traversed unless filtered or excluded by a filesystem boundary; scan errors mark totals incomplete.
## Filtering and Sorting
+13 -1
View File
@@ -4,7 +4,7 @@ A fast, concurrent, and feature-rich command-line tool for visualizing disk usag
## Features
- **Physical by Default**: Prioritizes **Disk Usage** (actual blocks consumed) as the default metric, essential for accurately seeing how much storage is actually gone.
- **Allocation by Default**: Prioritizes filesystem-reported allocated blocks; sparse files retain their separate apparent size, and hard-link allocation is counted once per scan. Reported allocation is not a promise of bytes freed by deletion on filesystems with shared extents or snapshots.
- **Fast & Concurrent**: Uses `rayon` to process directories in parallel, making it extremely fast on modern multi-core systems.
- **Rich Output**: Beautifully formatted tables with colors, distinguishing files and directories.
- **Apparent Size**: Toggle logical file length with `-a` or `--apparent`.
@@ -98,6 +98,7 @@ sized /path/to/directory
| `-f`, `--path-full` | Force absolute paths in headers | `sized -f` |
| `-i`, `--ignore` | Respect .gitignore files | `sized -i` |
| `-j`, `--threads` | Set number of threads | `sized -j 4` |
| `-x`, `--one-file-system` | Skip entries on other filesystems | `sized -x /` |
| `--format` | Output format (text, csv, json) | `sized --format json` |
| `--save` | Save output to file | `sized --save` |
| `-c`, `--compare` | Compare total vs. non-ignored files | `sized -i -c` |
@@ -137,6 +138,17 @@ autoload -Uz compinit && compinit
sized --completions fish > ~/.config/fish/completions/sized.fish
```
## Scan status and library API
Scans report missing or unreadable entries on stderr and return a nonzero status
when incomplete. JSON reports include `complete` so automation can distinguish
a partial report from a fully scanned tree. Symlinks are not followed.
Library users can call `scan_tree` with `ScanOptions` and a fresh `ScanControl`
for each request. Clones of the control expose progress and cancellation;
cancellation is cooperative between filesystem calls. `build_tree` remains as
a convenience helper, while `scan_tree` retains structured diagnostics.
## License
This project is licensed under the **GNU General Public License v3.0 (GPL-3.0)**.
+5 -5
View File
@@ -1,10 +1,10 @@
.ie \n(.g .ds Aq \(aq
.el .ds Aq '
.TH sized 1 "sized 1.0.0"
.TH sized 1 "sized 1.0.0"
.SH NAME
sized \- A modern, fast, and concurrent disk usage analyzer
.SH SYNOPSIS
\fBsized\fR [\fB\-v\fR|\fB\-\-version\fR] [\fB\-m\fR|\fB\-\-min\-size\fR] [\fB\-n\fR|\fB\-\-number\fR] [\fB\-\-sort\fR] [\fB\-a\fR|\fB\-\-apparent\fR] [\fB\-d\fR|\fB\-\-depth\fR] [\fB\-f\fR|\fB\-\-path\-full\fR] [\fB\-\-path\-relative\fR] [\fB\-j\fR|\fB\-\-threads\fR] [\fB\-\-completions\fR] [\fB\-i\fR|\fB\-\-ignore\fR] [\fB\-\-format\fR] [\fB\-\-save\fR] [\fB\-\-precision\fR] [\fB\-\-units\fR] [\fB\-c\fR|\fB\-\-compare\fR] [\fB\-h\fR|\fB\-\-help\fR] [\fIPATH\fR]
\fBsized\fR [\fB\-v\fR|\fB\-\-version\fR] [\fB\-m\fR|\fB\-\-min\-size\fR] [\fB\-n\fR|\fB\-\-number\fR] [\fB\-\-sort\fR] [\fB\-a\fR|\fB\-\-apparent\fR] [\fB\-d\fR|\fB\-\-depth\fR] [\fB\-f\fR|\fB\-\-path\-full\fR] [\fB\-\-path\-relative\fR] [\fB\-j\fR|\fB\-\-threads\fR] [\fB\-x\fR|\fB\-\-one\-file\-system\fR] [\fB\-\-completions\fR] [\fB\-i\fR|\fB\-\-ignore\fR] [\fB\-\-format\fR] [\fB\-\-save\fR] [\fB\-\-precision\fR] [\fB\-\-units\fR] [\fB\-c\fR|\fB\-\-compare\fR] [\fB\-h\fR|\fB\-\-help\fR] [\fIPATH\fR]
.SH DESCRIPTION
A modern, fast, and concurrent disk usage analyzer
.SH OPTIONS
@@ -36,6 +36,9 @@ Display paths relative to current directory (default)
\fB\-j\fR, \fB\-\-threads\fR \fI<THREADS>\fR
Number of threads to use (defaults to available logical CPUs)
.TP
\fB\-x\fR, \fB\-\-one\-file\-system\fR
Stay on the target\*(Aqs filesystem instead of descending into other mounts
.TP
\fB\-\-completions\fR \fI<COMPLETIONS>\fR
Generate shell completions
.br
@@ -76,6 +79,3 @@ Print help
Directory to analyze
.SH VERSION
v1.0.0
.SH COPYRIGHT
sized is licensed under the GNU General Public License v3.0 (GPL-3.0).
+59 -168
View File
@@ -6,15 +6,17 @@ use colored::*;
use comfy_table::presets::UTF8_FULL;
use comfy_table::{Attribute, Cell, CellAlignment, Color, ContentArrangement, Table};
use csv::WriterBuilder;
use ignore::WalkBuilder;
use rayon::prelude::*;
use serde::Serialize;
use std::cmp::Ordering;
use std::io::Write;
use std::os::unix::fs::MetadataExt;
use std::path::{Path, PathBuf};
use std::str::FromStr;
pub mod scan;
pub use scan::{
scan_tree, EntryType, FileIdentity, Node, ScanControl, ScanError, ScanIssue, ScanOptions,
ScanReport,
};
#[derive(Parser, Debug, Clone)]
#[command(author, version, about, long_about = None)]
#[command(disable_version_flag = true)]
@@ -61,6 +63,10 @@ pub struct Args {
#[arg(short = 'j', long = "threads")]
pub threads: Option<usize>,
/// Stay on the target's filesystem instead of descending into other mounts
#[arg(short = 'x', long)]
pub one_file_system: bool,
/// Generate shell completions
#[arg(long, value_enum)]
pub completions: Option<Shell>,
@@ -147,40 +153,14 @@ impl FromStr for SortDirection {
}
}
#[derive(Clone, Serialize, Debug)]
pub struct Node {
pub path: PathBuf,
pub size_bytes: u64,
pub blocks: u64,
pub size_bytes_filtered: u64,
pub blocks_filtered: u64,
pub entry_type: EntryType,
pub accessible: bool,
#[serde(skip)]
pub children: Vec<Node>,
}
#[derive(Clone, Serialize, Debug, PartialEq, Eq, PartialOrd, Ord)]
pub enum EntryType {
File,
Dir,
}
impl std::fmt::Display for EntryType {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
EntryType::File => write!(f, "File"),
EntryType::Dir => write!(f, "Dir"),
}
}
}
pub fn run(args: Args, mut writer: &mut dyn Write) {
/// Returns whether the scan completed without missing entries. The caller
/// chooses the exit status; partial reports remain available to the user.
pub fn run(args: Args, mut writer: &mut dyn Write) -> bool {
if let Some(shell) = args.completions {
let mut cmd = Args::command();
let name = cmd.get_name().to_string();
generate(shell, &mut cmd, name, &mut std::io::stdout());
return;
return true;
}
if let Some(out_dir) = args.generate_man_page {
@@ -193,23 +173,11 @@ pub fn run(args: Args, mut writer: &mut dyn Write) {
let file_path = out_dir.join("sized.1");
std::fs::write(&file_path, buffer).expect("Failed to write man page");
println!("Man page generated at {}", file_path.display());
return;
}
if let Some(threads) = args.threads {
rayon::ThreadPoolBuilder::new()
.num_threads(threads)
.build_global()
.ok(); // Ignore error if initialized called multiple times (e.g. in tests)
return true;
}
let target_path = args.path.clone();
if !target_path.exists() {
eprintln!("Error: Path '{}' does not exist.", target_path.display());
std::process::exit(1);
}
let min_bytes = if let Some(size_str) = &args.min_size {
match Byte::parse_str(size_str, true) {
Ok(byte) => byte.as_u64(),
@@ -234,134 +202,55 @@ pub fn run(args: Args, mut writer: &mut dyn Write) {
}
}
let root_node = build_tree(&target_path, args.ignore, args.compare);
let report = match scan_tree(
&target_path,
&ScanOptions {
respect_ignore: args.ignore,
compare_ignore: args.compare,
stay_on_filesystem: args.one_file_system,
threads: args.threads,
},
&ScanControl::default(),
) {
Ok(report) => report,
Err(e) => {
eprintln!("Error: {e}");
return false;
}
};
for issue in &report.issues {
eprintln!(
"Warning: incomplete scan at '{}': {}",
issue.path.display(),
issue.message
);
}
let root_node = report.root;
if root_node.size_bytes < min_bytes {
if args.format == OutputFormat::Text {
writeln!(writer, "Root directory is smaller than minimum size.").ok();
}
return;
return root_node.complete;
}
process_node_recursive(&mut writer, &root_node, 0, &args, min_bytes, &sort_criteria);
root_node.complete
}
pub fn build_tree(path: &Path, ignore: bool, compare: bool) -> Node {
let metadata = path.symlink_metadata();
if let Ok(meta) = metadata {
// Treat symlinks as files (nodes) but do not recurse
if meta.is_file() || meta.is_symlink() {
return Node {
path: path.to_path_buf(),
size_bytes: meta.len(),
blocks: meta.blocks(),
size_bytes_filtered: meta.len(), // Single file is its own filtered size for now
blocks_filtered: meta.blocks(),
entry_type: EntryType::File,
accessible: true,
children: vec![],
};
}
}
// Check if we can read the directory (handle permissions)
if let Err(e) = std::fs::read_dir(path) {
if e.kind() == std::io::ErrorKind::PermissionDenied {
// Return an "empty" directory node marked as inaccessible
let meta = path.symlink_metadata().ok();
let size = meta.as_ref().map(|m| m.len()).unwrap_or(0);
let blocks = meta.as_ref().map(|m| m.blocks()).unwrap_or(0);
return Node {
path: path.to_path_buf(),
size_bytes: size,
blocks,
size_bytes_filtered: size,
blocks_filtered: blocks,
entry_type: EntryType::Dir,
accessible: false,
children: vec![],
};
}
}
// Total walker (always everything if compare is true, else respects 'ignore' arg)
let total_ignore = if compare { false } else { ignore };
let walker_total = WalkBuilder::new(path)
.standard_filters(false)
.hidden(false)
.git_ignore(total_ignore)
.ignore(total_ignore)
.max_depth(Some(1))
.build();
let child_paths_total: Vec<PathBuf> = walker_total
.into_iter()
.filter_map(|e| e.ok())
.filter(|e| e.path() != path)
.map(|e| e.path().to_path_buf())
.collect();
// Filtered set (only if compare is true)
let non_ignored_set: std::collections::HashSet<PathBuf> = if compare {
let walker_filtered = WalkBuilder::new(path)
.standard_filters(false)
.hidden(false)
.git_ignore(true)
.ignore(true)
.max_depth(Some(1))
.build();
walker_filtered
.into_iter()
.filter_map(|e| e.ok())
.filter(|e| e.path() != path)
.map(|e| e.path().to_path_buf())
.collect()
} else {
std::collections::HashSet::new()
};
let children: Vec<Node> = child_paths_total
.par_iter()
.map(|p| build_tree(p, ignore, compare))
.collect();
let mut size_bytes = 0;
let mut blocks = 0;
let mut size_bytes_filtered = 0;
let mut blocks_filtered = 0;
for child in &children {
size_bytes += child.size_bytes;
blocks += child.blocks;
if compare {
if non_ignored_set.contains(&child.path) {
size_bytes_filtered += child.size_bytes_filtered;
blocks_filtered += child.blocks_filtered;
}
} else {
size_bytes_filtered += child.size_bytes_filtered;
blocks_filtered += child.blocks_filtered;
}
}
let (self_size, self_blocks) = path
.symlink_metadata()
.map(|m| (m.len(), m.blocks()))
.unwrap_or((0, 0));
Node {
path: path.to_path_buf(),
size_bytes: size_bytes + self_size,
blocks: blocks + self_blocks,
size_bytes_filtered: size_bytes_filtered + self_size, // self is always part of self
blocks_filtered: blocks_filtered + self_blocks,
entry_type: EntryType::Dir,
accessible: true,
children,
}
// Compatibility helper; callers needing diagnostics should use scan_tree.
scan_tree(
path,
&ScanOptions {
respect_ignore: ignore,
compare_ignore: compare,
..ScanOptions::default()
},
&ScanControl::default(),
)
.map(|report| report.root)
.unwrap_or_else(|_| Node::unavailable(path))
}
fn process_node_recursive(
@@ -372,8 +261,6 @@ fn process_node_recursive(
min_bytes: u64,
sort_criteria: &[(SortColumn, SortDirection)],
) {
if node.children.is_empty() && node.entry_type == EntryType::Dir {}
let mut display_children: Vec<&Node> = node
.children
.iter()
@@ -406,7 +293,7 @@ fn process_node_recursive(
print_output(
writer,
&node,
node,
&display_children,
args,
&relative_path,
@@ -735,7 +622,7 @@ fn print_text(
Cell::new("N/A")
.fg(Color::Red)
.set_alignment(CellAlignment::Right),
Cell::new("Access Denied")
Cell::new("Unavailable")
.fg(Color::Red)
.set_alignment(CellAlignment::Right),
];
@@ -794,6 +681,9 @@ fn print_json(writer: &mut dyn Write, parent_node: &Node, children: &[&Node], ar
"path": &c.path,
"entry_type": c.entry_type.to_string(),
"accessible": c.accessible,
"complete": c.complete,
"hard_link_duplicate": c.hard_link_duplicate,
"skipped_mount": c.skipped_mount,
"disk_usage": c.blocks * 512,
"blocks": c.blocks,
});
@@ -831,6 +721,7 @@ fn print_json(writer: &mut dyn Write, parent_node: &Node, children: &[&Node], ar
"total_blocks": parent_node.blocks,
"entries": entries_view,
"accessible": parent_node.accessible,
"complete": parent_node.complete,
});
if args.apparent {
+3 -1
View File
@@ -29,5 +29,7 @@ fn main() {
Box::new(std::io::stdout())
};
run(args, &mut writer);
if !run(args, &mut writer) {
std::process::exit(1);
}
}
+417
View File
@@ -0,0 +1,417 @@
//! Filesystem scanning, independent of CLI parsing and presentation.
//! Allocation is the filesystem's reported 512-byte blocks, not a prediction
//! of bytes freed by deletion (snapshots and shared extents can differ).
use ignore::WalkBuilder;
use rayon::prelude::*;
use serde::Serialize;
use std::{
collections::HashSet,
fs::{self, Metadata},
io,
os::unix::fs::MetadataExt,
path::{Path, PathBuf},
sync::{
atomic::{AtomicBool, AtomicU64, Ordering},
Arc, Mutex,
},
};
#[derive(Clone, Copy, Serialize, Debug, PartialEq, Eq)]
pub struct FileIdentity {
pub device: u64,
pub inode: u64,
pub links: u64,
}
#[derive(Clone, Serialize, Debug)]
pub struct Node {
pub path: PathBuf,
pub size_bytes: u64,
pub blocks: u64,
pub size_bytes_filtered: u64,
pub blocks_filtered: u64,
pub entry_type: EntryType,
pub accessible: bool,
/// False when this node or any descendant could not be scanned.
pub complete: bool,
pub symlink: bool,
pub skipped_mount: bool,
pub hard_link_duplicate: bool,
pub identity: Option<FileIdentity>,
#[serde(skip)]
pub children: Vec<Node>,
#[serde(skip)]
pub own_size_bytes: u64,
#[serde(skip)]
pub own_blocks: u64,
#[serde(skip)]
included_by_filter: bool,
}
#[derive(Clone, Serialize, Debug, PartialEq, Eq, PartialOrd, Ord)]
pub enum EntryType {
File,
Dir,
Special,
Unknown,
}
impl std::fmt::Display for EntryType {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{:?}", self)
}
}
impl Node {
pub(crate) fn unavailable(path: &Path) -> Self {
Self {
path: path.into(),
size_bytes: 0,
blocks: 0,
size_bytes_filtered: 0,
blocks_filtered: 0,
entry_type: EntryType::Unknown,
accessible: false,
complete: false,
symlink: false,
skipped_mount: false,
hard_link_duplicate: false,
identity: None,
children: Vec::new(),
own_size_bytes: 0,
own_blocks: 0,
included_by_filter: true,
}
}
fn from_metadata(path: &Path, meta: &Metadata) -> Self {
let mut node = Self::unavailable(path);
node.entry_type = if meta.is_dir() {
EntryType::Dir
} else if meta.is_file() || meta.is_symlink() {
EntryType::File
} else {
EntryType::Special
};
node.accessible = true;
node.complete = true;
node.symlink = meta.is_symlink();
node.identity = Some(FileIdentity {
device: meta.dev(),
inode: meta.ino(),
links: meta.nlink(),
});
node.own_size_bytes = meta.len();
node.own_blocks = meta.blocks();
node
}
}
#[derive(Clone, Debug, Default)]
pub struct ScanOptions {
pub respect_ignore: bool,
pub compare_ignore: bool,
pub stay_on_filesystem: bool,
/// None or zero selects Rayon's default worker count, scoped to this scan.
pub threads: Option<usize>,
}
/// Create a fresh control for each scan; clones share cancellation and progress.
#[derive(Clone, Debug, Default)]
pub struct ScanControl {
cancelled: Arc<AtomicBool>,
visited: Arc<AtomicU64>,
}
impl ScanControl {
pub fn cancel(&self) {
self.cancelled.store(true, Ordering::Relaxed);
}
pub fn is_cancelled(&self) -> bool {
self.cancelled.load(Ordering::Relaxed)
}
pub fn entries_scanned(&self) -> u64 {
self.visited.load(Ordering::Relaxed)
}
fn check(&self) -> Result<(), ScanError> {
if self.is_cancelled() {
Err(ScanError::Cancelled)
} else {
Ok(())
}
}
}
#[derive(Debug)]
pub struct ScanIssue {
pub path: PathBuf,
pub kind: io::ErrorKind,
pub message: String,
}
#[derive(Debug)]
pub struct ScanReport {
pub root: Node,
pub issues: Vec<ScanIssue>,
pub entries_scanned: u64,
}
#[derive(Debug)]
pub enum ScanError {
Root { path: PathBuf, source: io::Error },
WorkerPool(rayon::ThreadPoolBuildError),
Cancelled,
}
impl std::fmt::Display for ScanError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::Root { path, source } => {
write!(f, "cannot scan '{}': {}", path.display(), source)
}
Self::WorkerPool(e) => write!(f, "cannot start scan workers: {e}"),
Self::Cancelled => write!(f, "scan cancelled"),
}
}
}
impl std::error::Error for ScanError {}
/// The completed tree is published only after deterministic hard-link accounting.
/// Cancellation is cooperative between filesystem calls, not syscall preemption.
pub fn scan_tree(
path: &Path,
options: &ScanOptions,
control: &ScanControl,
) -> Result<ScanReport, ScanError> {
control.check()?;
let metadata = fs::symlink_metadata(path).map_err(|source| ScanError::Root {
path: path.into(),
source,
})?;
let pool = rayon::ThreadPoolBuilder::new()
.num_threads(options.threads.unwrap_or(0))
.build()
.map_err(ScanError::WorkerPool)?;
let context = Context {
options,
control,
root_device: metadata.dev(),
issues: Mutex::new(Vec::new()),
};
let mut root = pool.install(|| context.visit(path, Some(metadata)))?;
control.check()?;
reduce(
&mut root,
true,
&mut HashSet::new(),
&mut HashSet::new(),
control,
)?;
control.check()?;
let mut issues = context.issues.into_inner().unwrap();
issues.sort_by(|a, b| a.path.cmp(&b.path).then(a.message.cmp(&b.message)));
Ok(ScanReport {
root,
issues,
entries_scanned: control.entries_scanned(),
})
}
struct Context<'a> {
options: &'a ScanOptions,
control: &'a ScanControl,
root_device: u64,
issues: Mutex<Vec<ScanIssue>>,
}
impl Context<'_> {
fn issue(&self, path: &Path, kind: io::ErrorKind, message: String) {
self.issues.lock().unwrap().push(ScanIssue {
path: path.into(),
kind,
message,
});
}
fn visit(&self, path: &Path, metadata: Option<Metadata>) -> Result<Node, ScanError> {
self.control.check()?;
self.control.visited.fetch_add(1, Ordering::Relaxed);
let metadata = match metadata
.map(Ok)
.unwrap_or_else(|| fs::symlink_metadata(path))
{
Ok(meta) => meta,
Err(e) => {
self.issue(path, e.kind(), e.to_string());
return Ok(Node::unavailable(path));
}
};
let mut node = Node::from_metadata(path, &metadata);
if self.options.stay_on_filesystem && metadata.dev() != self.root_device {
node.skipped_mount = true;
node.own_size_bytes = 0;
node.own_blocks = 0;
return Ok(node);
}
if node.entry_type != EntryType::Dir {
return Ok(node);
}
if let Err(e) = fs::read_dir(path) {
self.issue(path, e.kind(), e.to_string());
node.accessible = false;
node.complete = false;
return Ok(node);
}
let total_ignore = self.options.respect_ignore && !self.options.compare_ignore;
let (paths, total_complete) = self.paths(path, total_ignore)?;
let (filtered, filtered_complete) = if self.options.compare_ignore {
let (paths, complete) = self.paths(path, true)?;
(paths.into_iter().collect::<HashSet<_>>(), complete)
} else {
(HashSet::new(), true)
};
node.complete = total_complete && filtered_complete;
node.children = paths
.par_iter()
.map(|child| {
let mut node = self.visit(child, None)?;
node.included_by_filter = !self.options.compare_ignore || filtered.contains(child);
Ok(node)
})
.collect::<Result<Vec<_>, ScanError>>()?;
Ok(node)
}
fn paths(&self, path: &Path, respect_ignore: bool) -> Result<(Vec<PathBuf>, bool), ScanError> {
let walker = WalkBuilder::new(path)
.standard_filters(false)
.hidden(false)
.parents(respect_ignore)
.git_ignore(respect_ignore)
.ignore(respect_ignore)
.max_depth(Some(1))
.build();
let mut paths = Vec::new();
let mut complete = true;
for entry in walker {
self.control.check()?;
match entry {
Ok(entry) if entry.path() != path => paths.push(entry.into_path()),
Ok(_) => {}
Err(e) => {
complete = false;
self.issue(
path,
e.io_error().map_or(io::ErrorKind::Other, |e| e.kind()),
e.to_string(),
);
}
}
}
// Parallel collection preserves this order; ownership of a shared inode
// is then stable across worker counts and repeated scans.
paths.sort();
Ok((paths, complete))
}
}
fn reduce(
node: &mut Node,
included: bool,
all_seen: &mut HashSet<(u64, u64)>,
filtered_seen: &mut HashSet<(u64, u64)>,
control: &ScanControl,
) -> Result<(), ScanError> {
control.check()?;
node.size_bytes = node.own_size_bytes;
node.blocks = node.own_blocks;
node.size_bytes_filtered = if included { node.own_size_bytes } else { 0 };
node.blocks_filtered = if included { node.own_blocks } else { 0 };
if node.entry_type != EntryType::Dir {
if let Some(id) = node.identity.filter(|id| id.links > 1) {
let key = (id.device, id.inode);
if !all_seen.insert(key) {
node.blocks = 0;
node.hard_link_duplicate = true;
}
if included && !filtered_seen.insert(key) {
node.blocks_filtered = 0;
}
}
}
for child in &mut node.children {
reduce(
child,
included && child.included_by_filter,
all_seen,
filtered_seen,
control,
)?;
node.size_bytes += child.size_bytes;
node.blocks += child.blocks;
node.size_bytes_filtered += child.size_bytes_filtered;
node.blocks_filtered += child.blocks_filtered;
node.complete &= child.complete;
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use tempfile::tempdir;
#[test]
fn foreign_device_entry_is_marked_without_scanning_its_contents() {
let dir = tempdir().unwrap();
fs::write(dir.path().join("content"), b"not visited").unwrap();
let meta = fs::symlink_metadata(dir.path()).unwrap();
let control = ScanControl::default();
let options = ScanOptions {
stay_on_filesystem: true,
..ScanOptions::default()
};
// Inject the parent scan's different device; this exercises the boundary
// policy without creating mounts or requiring privileged test setup.
let context = Context {
options: &options,
control: &control,
root_device: meta.dev().wrapping_add(1),
issues: Mutex::new(Vec::new()),
};
let mut node = context.visit(dir.path(), Some(meta)).unwrap();
reduce(
&mut node,
true,
&mut HashSet::new(),
&mut HashSet::new(),
&control,
)
.unwrap();
assert!(node.skipped_mount);
assert!(node.children.is_empty());
assert_eq!(node.blocks, 0);
assert_eq!(control.entries_scanned(), 1);
assert!(node.complete);
}
#[test]
fn vanished_child_is_reported_without_fabricating_accessibility() {
let dir = tempdir().unwrap();
let path = dir.path().join("vanished-after-discovery");
let options = ScanOptions::default();
let control = ScanControl::default();
let context = Context {
options: &options,
control: &control,
root_device: fs::metadata(dir.path()).unwrap().dev(),
issues: Mutex::new(Vec::new()),
};
let node = context.visit(&path, None).unwrap();
assert!(!node.accessible);
assert!(!node.complete);
let issues = context.issues.into_inner().unwrap();
assert_eq!(issues[0].path, path);
assert_eq!(issues[0].kind, io::ErrorKind::NotFound);
}
}
+52
View File
@@ -45,3 +45,55 @@ fn test_json_output() {
.stdout(predicate::str::contains("\"entry_type\":\"File\""))
.stdout(predicate::str::contains("\"path\":"));
}
#[test]
fn missing_path_fails_without_a_success_report() {
let dir = TempDir::new().unwrap();
Command::new(env!("CARGO_BIN_EXE_sized"))
.arg(dir.path().join("missing"))
.args(["--format", "json"])
.assert()
.failure()
.stdout(predicate::str::is_empty())
.stderr(predicate::str::contains("cannot scan"));
}
#[cfg(unix)]
#[test]
fn partial_scan_has_nonzero_status_and_explicit_json_flag() {
use std::{fs, os::unix::fs::PermissionsExt};
let dir = TempDir::new().unwrap();
let blocked = dir.path().join("blocked");
fs::create_dir(&blocked).unwrap();
fs::set_permissions(&blocked, fs::Permissions::from_mode(0o000)).unwrap();
let result = std::panic::catch_unwind(|| {
assert!(
fs::read_dir(&blocked).is_err(),
"permission tests require an unprivileged user"
);
Command::new(env!("CARGO_BIN_EXE_sized"))
.arg(dir.path())
.args(["--format", "json"])
.assert()
.failure()
.stdout(predicate::str::contains("\"complete\":false"))
.stderr(predicate::str::contains("incomplete scan"));
});
fs::set_permissions(blocked, fs::Permissions::from_mode(0o700)).unwrap();
result.unwrap();
}
#[cfg(unix)]
#[test]
fn dangling_symlink_is_a_valid_scan_target() {
use std::os::unix::fs::symlink;
let dir = TempDir::new().unwrap();
let path = dir.path().join("dangling");
symlink(dir.path().join("missing"), &path).unwrap();
Command::new(env!("CARGO_BIN_EXE_sized"))
.arg(path)
.args(["--format", "json"])
.assert()
.success()
.stdout(predicate::str::contains("\"complete\":true"));
}
+254
View File
@@ -0,0 +1,254 @@
#![cfg(unix)]
use sized::{build_tree, scan_tree, EntryType, ScanControl, ScanError, ScanOptions};
use std::{
fs,
io::Write,
os::unix::fs::{symlink, MetadataExt, PermissionsExt},
};
use tempfile::tempdir;
fn scan(path: &std::path::Path) -> sized::ScanReport {
scan_tree(path, &ScanOptions::default(), &ScanControl::default()).unwrap()
}
#[test]
fn hard_links_count_allocation_once_but_keep_apparent_sizes() {
let dir = tempdir().unwrap();
let original = dir.path().join("a-original");
fs::write(&original, [9; 65536]).unwrap();
fs::hard_link(&original, dir.path().join("z-alias")).unwrap();
let expected =
fs::metadata(&original).unwrap().blocks() + fs::metadata(dir.path()).unwrap().blocks();
for threads in [1, 4] {
let report = scan_tree(
dir.path(),
&ScanOptions {
threads: Some(threads),
..ScanOptions::default()
},
&ScanControl::default(),
)
.unwrap();
assert_eq!(report.root.blocks, expected);
assert_eq!(report.root.children[0].size_bytes, 65536);
assert_eq!(report.root.children[1].size_bytes, 65536);
assert!(!report.root.children[0].hard_link_duplicate);
assert!(report.root.children[1].hard_link_duplicate);
assert_eq!(report.root.children[1].blocks, 0);
assert_eq!(report.entries_scanned, 3);
assert!(report.root.complete);
}
}
#[test]
fn ignored_hard_link_does_not_steal_filtered_allocation() {
let dir = tempdir().unwrap();
fs::write(dir.path().join(".ignore"), "a-ignored\n").unwrap();
let original = dir.path().join("a-ignored");
fs::write(&original, [9; 65536]).unwrap();
fs::hard_link(&original, dir.path().join("z-included")).unwrap();
let report = scan_tree(
dir.path(),
&ScanOptions {
compare_ignore: true,
..ScanOptions::default()
},
&ScanControl::default(),
)
.unwrap();
let expected = fs::metadata(&original).unwrap().blocks()
+ fs::metadata(dir.path()).unwrap().blocks()
+ fs::metadata(dir.path().join(".ignore")).unwrap().blocks();
assert_eq!(report.root.blocks, expected);
assert_eq!(report.root.blocks_filtered, expected);
assert_eq!(report.root.children.last().unwrap().blocks, 0);
assert!(report.root.children.last().unwrap().blocks_filtered > 0);
}
#[test]
fn hidden_files_are_counted_without_ignore_filtering() {
let dir = tempdir().unwrap();
fs::write(dir.path().join(".hidden"), b"hidden").unwrap();
fs::write(dir.path().join(".ignore"), "build\n").unwrap();
fs::write(dir.path().join("build"), b"build").unwrap();
assert_eq!(scan(dir.path()).root.children.len(), 3);
let report = scan_tree(
dir.path(),
&ScanOptions {
respect_ignore: true,
..ScanOptions::default()
},
&ScanControl::default(),
)
.unwrap();
assert_eq!(report.root.children.len(), 2);
assert!(report
.root
.children
.iter()
.any(|n| n.path.ends_with(".hidden")));
}
#[test]
fn ignore_rules_are_inherited_by_nested_directories() {
let dir = tempdir().unwrap();
fs::write(dir.path().join(".ignore"), "*.tmp\n").unwrap();
let nested = dir.path().join("nested");
fs::create_dir(&nested).unwrap();
fs::write(nested.join("skip.tmp"), b"ignored").unwrap();
fs::write(nested.join("keep.bin"), b"included").unwrap();
let report = scan_tree(
dir.path(),
&ScanOptions {
respect_ignore: true,
..ScanOptions::default()
},
&ScanControl::default(),
)
.unwrap();
let children = &report
.root
.children
.iter()
.find(|n| n.path == nested)
.unwrap()
.children;
assert_eq!(children.len(), 1);
assert!(children[0].path.ends_with("keep.bin"));
let compared = scan_tree(
dir.path(),
&ScanOptions {
compare_ignore: true,
..ScanOptions::default()
},
&ScanControl::default(),
)
.unwrap();
assert_eq!(compared.root.size_bytes_filtered, report.root.size_bytes);
assert_eq!(compared.root.blocks_filtered, report.root.blocks);
}
#[test]
fn sparse_file_reports_blocks_and_logical_length_separately() {
let dir = tempdir().unwrap();
let path = dir.path().join("sparse");
let mut file = fs::File::create(&path).unwrap();
file.write_all(&[7; 4096]).unwrap();
file.set_len(128 * 1024 * 1024).unwrap();
file.sync_all().unwrap();
let node = scan(&path).root;
assert_eq!(node.size_bytes, 128 * 1024 * 1024);
assert_eq!(node.blocks, fs::metadata(path).unwrap().blocks());
assert!(node.blocks * 512 < node.size_bytes);
}
#[test]
fn symlinks_including_loop_and_dangling_link_are_leaves() {
let dir = tempdir().unwrap();
symlink(dir.path(), dir.path().join("loop")).unwrap();
symlink(dir.path().join("missing"), dir.path().join("dangling")).unwrap();
let report = scan(dir.path());
assert_eq!(report.root.children.len(), 2);
for node in report.root.children {
assert!(node.symlink);
assert!(node.children.is_empty());
assert!(node.complete);
}
}
#[test]
fn missing_root_returns_error_and_legacy_helper_marks_unavailable() {
let dir = tempdir().unwrap();
let path = dir.path().join("vanished");
assert!(
matches!(scan_tree(&path, &ScanOptions::default(), &ScanControl::default()),
Err(ScanError::Root { source, .. }) if source.kind() == std::io::ErrorKind::NotFound)
);
let node = build_tree(&path, false, false);
assert_eq!(node.entry_type, EntryType::Unknown);
assert!(!node.accessible);
assert!(!node.complete);
}
#[test]
fn permission_error_propagates_incomplete_status_to_root() {
let dir = tempdir().unwrap();
let blocked = dir.path().join("blocked");
fs::create_dir(&blocked).unwrap();
fs::write(blocked.join("unreadable"), b"data").unwrap();
fs::set_permissions(&blocked, fs::Permissions::from_mode(0o000)).unwrap();
let result = std::panic::catch_unwind(|| {
assert!(
fs::read_dir(&blocked).is_err(),
"permission tests require an unprivileged user"
);
let report = scan(dir.path());
assert!(report.root.accessible);
assert!(!report.root.complete);
assert!(!report.root.children[0].accessible);
assert_eq!(report.issues.len(), 1);
assert_eq!(report.issues[0].path, blocked);
assert_eq!(report.issues[0].kind, std::io::ErrorKind::PermissionDenied);
});
fs::set_permissions(blocked, fs::Permissions::from_mode(0o700)).unwrap();
result.unwrap();
}
#[test]
fn cancelled_scan_never_returns_completed_tree() {
let dir = tempdir().unwrap();
let control = ScanControl::default();
control.cancel();
assert!(matches!(
scan_tree(dir.path(), &ScanOptions::default(), &control),
Err(ScanError::Cancelled)
));
assert_eq!(control.entries_scanned(), 0);
}
#[test]
fn cancellation_and_progress_work_during_scan() {
let dir = tempdir().unwrap();
for i in 0..2000 {
fs::write(dir.path().join(format!("file-{i}")), b"x").unwrap();
}
let control = ScanControl::default();
let watcher = control.clone();
let cancel = std::thread::spawn(move || {
let deadline = std::time::Instant::now() + std::time::Duration::from_secs(5);
while watcher.entries_scanned() < 10 {
assert!(
std::time::Instant::now() < deadline,
"scan made no progress"
);
std::thread::yield_now();
}
watcher.cancel();
});
let result = scan_tree(
dir.path(),
&ScanOptions {
threads: Some(1),
..ScanOptions::default()
},
&control,
);
cancel.join().unwrap();
assert!(matches!(result, Err(ScanError::Cancelled)));
assert!(control.entries_scanned() >= 10);
}
#[test]
fn special_file_is_not_treated_as_directory() {
let dir = tempdir().unwrap();
let path = dir.path().join("fifo");
assert!(std::process::Command::new("mkfifo")
.arg(&path)
.status()
.unwrap()
.success());
let node = scan(&path).root;
assert_eq!(node.entry_type, EntryType::Special);
assert!(node.children.is_empty());
assert!(node.complete);
}