Skip to content

fix: garbled zip entry names in non-UTF-8 encodings - #1089

Open
Mathjk wants to merge 2 commits into
ouch-org:mainfrom
Mathjk:fix/non-utf8-filenames
Open

Mathjk wants to merge 2 commits into
ouch-org:mainfrom
Mathjk:fix/non-utf8-filenames

Conversation

@Mathjk

@Mathjk Mathjk commented Sep 29, 2026

Copy link
Copy Markdown

Fixes #691.

Problem

When a zip's entry names aren't UTF-8 and the archive's UTF-8 flag (bit 11) is unset, the zip crate falls back to CP437 — so archives written on systems using GBK/Big5/Shift_JIS/windows-125x, or UTF-8 names written without the flag, come out as mojibake that can't even be fixed afterwards with convmv/iconv (the decoded text is already wrong). unzip/7zz recover such archives fine.

Changes

  • New --encoding <ENCODING> flag on ouch decompress and ouch list, following unzip -O precedent. Accepts WHATWG encoding labels (gbk, big5, shift_jis, windows-1251, ...) plus common code-page aliases (cp936, windows-1252, cp65001). Unknown labels get a clean Unknown encoding error. Zip-only; documented in --help.
  • Unconditional safe fix: unflagged names that are already valid UTF-8 now decode as UTF-8 (previously CP437-garbled). Names flagged UTF-8 are untouched (spec). Everything else still falls back to CP437 — zero regression risk for spec-compliant archives.
  • enclosed_name/mangled_name reimplemented via typed-path (the same crate zip uses internally) so sanitization runs on the decoded name — this also closes a subtle bug where a multibyte trail byte 0x5C (e.g. in GBK 说) could be misread as a path separator.
  • Directory detection now uses file.is_dir() instead of checking for a trailing /.

Testing

  • zip -r archive with raw GBK (CP936) filename bytes → before: test-name-╒Γ╩╟╥╗╕÷▓Γ╩╘╬─╝■.txt; after --encoding gbk: test-name-这是一个测试文件.txt ✓
  • Unflagged-but-UTF-8 name (Schwarz-weiß.txt) now recovers with no flag (issue's German-umlaut case) ✓
  • --encoding not-a-charset → clean error ✓
  • 6 new unit tests incl. an in-test GBK-zip fixture; full suite green (cargo test --profile fast: 59 unit / 57 integration / 13 ui / 3 mime); cargo clippy --all-targets clean incl. --no-default-features; nightly cargo fmt applied.

Discussion point

Auto-detection (e.g. chardetng) was considered instead of a flag but rejected: it misdecodes genuine CP437 names (verified: CP437 Grüße → Shift_JIS Gr≪e) and would silently change existing behavior. Happy to add detection behind a flag value (--encoding auto) if you'd like.

New deps: encoding_rs (+6 small transitive crates), typed-path (already in tree via zip), crc32fast (dev-dep for test fixtures).

Zip entry names that do not carry the UTF-8 flag were always decoded as
CP437 (the ZIP spec default), which produces mojibake for archives
created by tools that store names in UTF-8 without the flag, or in a
legacy code page such as GBK (CP936), Big5 or Shift_JIS (issue ouch-org#691).

- Add `--encoding <ENCODING>` to `decompress` and `list`. It accepts
  WHATWG encoding labels (gbk, big5, shift_jis, windows-1251, ...) and
  the familiar code page names that WHATWG does not label (cp936,
  windows-1252, ...), resolved with the new encoding_rs dependency.
- Entry names flagged as UTF-8 are unaffected, per the ZIP spec.
- Without `--encoding`, unflagged names that are valid UTF-8 are now
  decoded as UTF-8 instead of CP437, recovering archives written by
  tools that store UTF-8 names but forget to set the flag.
- Otherwise the spec-mandated CP437 decode is kept as the fallback.

Path safety checks (enclosed_name/mangled_name) now run on the decoded
name; they are reimplemented on top of typed_path, the same crate the
zip crate uses, so the semantics are unchanged. Decoding before
sanitizing also fixes multi-byte names whose trail bytes happen to be
a backslash byte, which used to be treated as path separators.
@valoq

valoq commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator
  • zip.rs:358-371: absolute entry paths (/abs/evil.txt) are extracted under the output dir instead of rejected but the doc comment says it matches ZipFile::enclosed_name
  • decompress.rs:224: --encoding is silently ignored for non-zip archives, please add a warning.
  • zip.rs:350: decode errors are dropped and a wrong --encoding gives U+FFFD in names with no warning.

@Mathjk

Mathjk commented Oct 8, 2026

Copy link
Copy Markdown
Author

Addressed in cebf2e8. enclosed_name now rejects absolute entry paths outright — stricter than zip 8.6's own impl, which silently strips a leading root (its doc says absolute paths are rejected; the impl doesn't quite). --encoding on non-zip archives now warns in both decompress and list. Decode errors under --encoding warn when replacement chars are produced. Tests added for all three.

@valoq

valoq commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Files with a name whose last byte is 0x5C is now extracted as an empty directory and the file content is lost

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ouch breaks the file name when the encode is not utf-8

2 participants