Use byte widths for fixed-record MIME databases - #65
Closed
OskarEichler wants to merge 1 commit into
Closed
Conversation
Merged
SamSaffron
added a commit
that referenced
this pull request
Sep 2, 2026
Require Ruby 3.3 and refresh the bundled MIME database from mime-types-data 3.2026.0701. Use byte offsets for Unicode-safe database lookups, avoid duplicate lowercase misses, and preserve extension priorities when rebuilding the database. Update CI to cover current Ruby implementations and runners. closes #64, #65, #66
Member
|
sorry addressed in my big pr |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Measure database row sizes and generated column padding in bytes. Positional reads use byte offsets, but String#length counts characters; custom Unicode extensions could make later records unreadable.
Reproduction
Write a UTF-8 custom extension database with these equal-byte-width rows:
RandomAccessDb.new(path, 0).lookup('ê') returned nil. It now returns text/two. Three focused lookup/miss assertions pass with native pread and again with the fallback seek/read implementation selected in a disposable process. A generator check confirms mixed ASCII/Unicode columns produce equal byte widths. No database regeneration is included.
Verification
Compatibility
No intended breaking change. ASCII database bytes and lookup results are unchanged; Unicode custom databases must use fixed byte widths, as required by positional reads. The fallback check is simulated on macOS, not a Windows CI claim.