Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions .github/workflows/release-gate.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,11 @@ concurrency:

jobs:
release-gate:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
runs-on: ${{ matrix.os }}
timeout-minutes: 10
steps:
- name: Check out repository
Expand All @@ -41,7 +45,7 @@ jobs:
run: python -W error::ResourceWarning -m unittest discover -s tests -v

- name: Compile Python sources
run: python -m py_compile run_osintai.py src/osintai/*.py tests/*.py
run: python -m compileall -q src tests run_osintai.py

- name: Run security analysis
run: bandit -q -r src run_osintai.py -ll -ii
Expand All @@ -51,5 +55,5 @@ jobs:

- name: Verify release identity
run: |
python run_osintai.py --version | grep -Fx "OSINTai 4.0.0"
python run_osintai.py --help >/dev/null
python -c "import sys; sys.path.insert(0, 'src'); from osintai import __version__; assert f'# OSINTai v{__version__} ' in open('README.md', encoding='utf-8').read()"
python run_osintai.py --help
4 changes: 2 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,7 @@ id_ed25519
*.ovpn

# Local OSINT inputs/outputs may contain PII, targets, proxies, or findings
data/runs/
data/cache/
/data/
cases/
/OSINTAI_SYNTHESIS/
/TODO
Expand Down Expand Up @@ -62,6 +61,7 @@ build/
.pytest_cache/
.mypy_cache/
.tox/
.test-tmp/

# IDEs
.vscode/
Expand Down
62 changes: 62 additions & 0 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# OSINTai 4.2.0 offline benchmarks

Measured on 2026-09-10 using Python 3.12.13, Darwin
arm64. These are single-run observations, not statistical guarantees.
No page fetching or model calls were performed. Source captures and existing reports
were preserved; the benchmark produced new extraction checkpoints and report bundles.

## Saved crawls

Default analysis settings: 200,000 characters per page; 16 MB text cache per reader;
100,000 co-occurrence candidate-pair budget; 30-second page and 120-second stage deadlines.
Timing includes source hashing, worker startup, analysis, report validation, and publication.

| Saved run | Cache | Pages | Elapsed | Cache hits/misses | Parent peak RSS | Maximum child RSS |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| 20260909_212714 | cold | 152 | 18.133 s | 0/152 | 91.9 MiB | 84.4 MiB |
| 20260909_212714 | warm | 152 | 1.644 s | 152/0 | 86.8 MiB | 84.4 MiB |
| 20260717_144145 | cold | 159 | 18.923 s | 2/157 | 85.5 MiB | 92.6 MiB |

The 159-page cold run reused two identical page contents within the same invocation.
RSS is reported separately for the parent and the largest individual child; summing the
columns is not a measurement of simultaneous process-tree memory. Warm means extraction
checkpoints were reused, not that every filesystem or OS cache was controlled.

Both crawls explicitly report partial analytical coverage. The 152-page crawl excluded
92,369 potential pairs on pages exceeding the 60-identifier pairing threshold. The
159-page crawl excluded 161,328 such pairs, omitted 25,125 low-ranked correlation rows
at the output cap, and recorded 46 temporal parsing errors from source date values.
Neither run exhausted its candidate-pair budget. Publication succeeded and these limits
and errors remain visible in each bundle's summary and report.

## Boilerplate fixture

The synthetic fixture has 10,000 pages, four shared footer identifiers, and one
unique email/handle pair per page. All four footer identifiers are excluded from pairing.

- Entity indexing: **0.053 seconds**.
- Correlation: **0.079 seconds**.
- Examined co-occurrence pairs: **10,000**.
- Pairs omitted by budget: **0**.
- Parent peak RSS: **60.8 MiB**.

## Reproduction

Run from the repository using its Python environment. Benchmark memory collection uses
`resource`, available on macOS/Linux; the unit-test CI matrix also covers Windows.

```bash
python tests/benchmark_analysis.py RUN_ID --output .test-tmp/benchmark-cold.json
python tests/benchmark_analysis.py RUN_ID --output .test-tmp/benchmark-warm.json
python tests/benchmark_correlation.py --pages 10000 --pair-budget 100000
```

A run is cold only if no matching extractor-version/content checkpoints exist. No
benchmark deletes prior checkpoints or source material to manufacture a cold run.

## Verification

All **116 offline tests** passed locally with `ResourceWarning` treated as an error.
Correctness lint, Bandit at the release gate's severity/confidence thresholds, compilation,
CLI version/help checks, and whitespace checks passed. Linux/Windows CI execution remains
pending; configuring that matrix does not establish a passing result on those platforms.
95 changes: 95 additions & 0 deletions ENHANCEMENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# OSINTai enhancements

## Integrated in 4.2.0

All seven enhancement areas from the 4.1.0 follow-up list are implemented:

1. **Process-isolated deadlines.** Live HTML parsing, indicator extraction, hunt matching,
and Simhash run in disposable spawn workers. Saved-page extraction, entity indexing,
deterministic/model stages, and training export have deadlines. CPU timeouts, crashed
children, cancellation, and large IPC results have offline tests. Workers are terminated
and reaped. Raw HTML is retained before parsing, and failures are recorded. CLI controls:
`--analysis-page-timeout` (30 seconds), `--analysis-stage-timeout` (120 seconds).
2. **Content-addressed extraction checkpoints.** `.extraction_cache/` keys include scanned
text SHA-256, extractor implementation hash, character limit, and truncation state.
Atomic JSON checkpoints carry schema/checksum validation and per-kind coverage counts.
Repeated pages and resumed analyses reuse valid results. Cache writes failing does not
discard successful extraction. Secret values and arbitrary JWT claims are excluded;
secret-shaped values are also removed from cached indicator fields, including values
beyond the findings cap. Reports reconstruct findings from counts and decode validity.
3. **Adversarial regex audit.** Email, ASCII-domain, credential, and JWT extraction consume
maximal tokens or lines before bounded validation. Oversized malformed candidates cannot
restart an unbounded suffix search. Unicode scanning remains linear. External-deadline
tests cover repeated punctuation, failed domain/email suffixes, JWT-like strings, and
long credential lines. Fixed detector limits and their omissions are reported.
4. **Model quality and bounded retry.** Results distinguish `ok`, `empty`, `invalid`,
`missing`, `timed_out`, `error`, and `skipped`. Page attribution comes from saved file
identity rather than model assertions. Manifests record model names, validated page
counts, and optional stage response counts. `--retry-model RUN_ID --retry-limit 20
--retry-timeout 60` retries unsuccessful/skipped saved pages through local Ollama,
without fetching pages or overwriting original analyses. Later offline runs use the
published `model_retry_latest.json` overlay.
5. **Hunt offsets and URL provenance.** Unicode expansion maps matches back to original
start/end offsets. URL detection uses original text and rejects URLs crossing a snippet
boundary. Indicator provenance identifies `html_attribute`, `page_prose`, and
`html_source`; resolved relative HTML links are included.
6. **Bounded correlation and text caching.** `--correlation-pair-budget` defaults to
100,000 examined co-occurrence pairs. Omitted pairs, oversized-page exclusions, and
output truncation are explicit. Source membership uses sets, and identity evidence
avoids repeatedly scanning a common handle's entire source list. Each text reader has
a `--text-cache-bytes` LRU budget (16 MB default), accounting for string and entry
overhead. Cache peaks, evictions, missing files, and truncation are reported.
7. **Transactional analysis publication.** A hidden incomplete directory is populated
and validated before an atomic rename publishes an immutable `analysis_*/` bundle.
The bundle includes the human report, JSON/JSONL results, summary, manifest, and optional
training export. `analysis_latest.json` advances only after successful publication.
Attempt status records running/failed/completed state, and manifests retain full source
hashes and options. Source changes during analysis prevent publication. File contents
are flushed; POSIX directory metadata is fsynced. Original captures and prior bundles
remain intact. Completed publication may still report partial analytical coverage.

The existing 4.1.0 fixes are retained: Unicode scanning, offline recovery, bounded text
reads, failure isolation, clean interruption, HTML text boundaries, complete summaries,
empty artifacts, numeric validation, and explicit profile overrides.

## Verification and performance

See `tests/test_enhancements.py` for deterministic offline acceptance checks and
`tests/benchmark_analysis.py` for saved-crawl runtime and memory measurements. CI now
runs the release gate on macOS, Linux, and Windows. Local verification was performed on
macOS with Python 3.12; the Linux/Windows matrix needs to execute in CI.

Measured results are recorded in `BENCHMARKS.md`. RSS figures separate parent memory
from the maximum child RSS; they are not a measured concurrent process-tree total.

Scanning is O(N) for bounded token validators and fixed extraction caps. Entity indexing
is expected O(S), with S source/indicator observations. Correlation generation is bounded
by O(S + B), followed by O(K log K) ordering of K candidates, where B is the pair budget.
Core retained analysis memory is O(S + E + B + C + R + L): E entities, C the text-cache budget,
R retained result rows, and L the per-job input bound. Disposable workers add serialization copies and startup cost.
Hunt searches cost O(TN) for T terms, with at most 500 reported hits; Unicode offset mapping
uses O(N) space only when lowercasing expands characters. Full source hashing streams
files with a 1 MB buffer. Thread timeouts and unbounded caches were rejected because they
cannot provide containment and predictable memory use.

## Recommended next work

1. **Public-suffix-aware domain families.** `_registrable()` still uses the last two
labels; names beneath suffixes such as `co.uk` can be incorrectly grouped. Bundle a
versioned Public Suffix List for deterministic offline grouping and ownership caveats.
2. **Cache retention and quotas.** Add explicit age/size-based pruning for extraction
checkpoints and old report/retry bundles, preserving referenced evidence and active
attempts. Current cache memory is bounded, but disk retention is intentionally additive.
3. **Typed streaming artifact ingestion.** Stream large JSONL files and validate record
schemas with line-level corruption counts. Current loaders materialize metadata and
silently skip malformed JSON lines; text-cache limits do not bound total index memory.
4. **Reduce process startup overhead.** Benchmark a supervised, recyclable worker design
that preserves hard per-job termination and isolation. The current spawn-per-job design
is portable and simple, but cold runs pay startup cost for every distinct page.
5. **More precise reproducibility and memory telemetry.** Inject a reference clock for
temporal analysis and generated timestamps; measure concurrent process-tree RSS, and
add Windows benchmark memory collection. Validate filesystem crash durability on the
supported filesystems; Windows does not provide POSIX directory-fsync semantics here.
6. **Raw-only parsing recovery.** Offer an explicit offline command to reparse preserved
HTML from live extraction failures. Current `--analyze-only` consumes saved page text
and reports raw-only failures rather than reconstructing missing text automatically.
104 changes: 100 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
<img src="bd800949-d4d1-44ce-849e-ba40837590bc.png" alt="OSINTai Logo" width="100%">

# OSINTai v4 - Advanced Local-First OSINT Web Crawler
# OSINTai v4.2.0 - Advanced Local-First OSINT Web Crawler

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
Expand Down Expand Up @@ -253,7 +253,7 @@ usage: run_osintai.py [-h] [--seed SEED] [--depth DEPTH] [--max MAX]
[--no-ollama] [--hunt HUNT] [--hunt-max HUNT_MAX]
[--run-id RUN_ID]

OSINTai v4 (async crawling and analysis)
OSINTai 4.2.0 (async crawling and analysis)

required arguments:
--seed SEED Seed URL (or use seed_urls.txt file)
Expand Down Expand Up @@ -290,6 +290,80 @@ optional deeper analysis (off by default):

### Analysis Layer

**Recover a completed crawl without fetching pages again:**

```bash
python run_osintai.py --analyze-only 20260909_212714
```

This mode is fully offline and does not contact Ollama. Each invocation writes to a new
`data/runs/RUN_ID/reanalysis_*/` directory, preserving the source crawl and earlier reports.
It can also use `--evaluate` and `--training-export`; model-assisted modes are rejected.
It uses existing saved page analyses and the latest explicit model-retry overlay. It does
not regenerate missing model responses.

### Analysis reliability and recovery (4.2)

Live HTML parsing and saved-page extraction run in disposable worker processes with a
30-second deadline. CPU analysis stages, including indexing and dataset export, have a
120-second deadline. A failed worker is terminated and reaped; the remaining analysis
records partial coverage and continues. Raw HTML is saved before live parsing.

```bash
python run_osintai.py --analyze-only RUN_ID \
--analysis-page-timeout 30 --analysis-stage-timeout 120 \
--analysis-max-chars 200000 --text-cache-bytes 16000000 \
--correlation-pair-budget 100000
```

Successful page extraction is cached under `RUN_ID/.extraction_cache/`. Keys include the
scanned-text SHA-256, extractor implementation version, character limit, and truncation
state. Repeated or interrupted runs reuse valid results. Checkpoints retain secret counts,
entropy counts, and JWT decode validity; matched secret values and arbitrary JWT claims
are excluded. Original page captures remain unchanged. Cache hits, invalid entries,
extraction failures, text coverage, and per-kind indicator omissions are reported.

Text readers use an LRU cache with a byte budget that accounts for strings and entry
overhead. Correlation examines at most the configured number of co-occurrence pairs.
Its stage statistics distinguish budget omissions, oversized-page omissions, and output
row truncation. Common footer entities are excluded from pairing and indexed with sets.
Hunt offsets refer to original text even when Unicode lowercasing expands characters;
URLs crossing a snippet boundary are excluded. `url_provenance` distinguishes HTML
links, prose URLs, and URLs found elsewhere in HTML source.

Model results distinguish `ok`, `empty`, `invalid`, `missing`, `timed_out`, `error`, and
`skipped`, with per-model counts in the manifest. Explicitly retry unsuccessful or skipped
saved-page model analyses using local Ollama:

```bash
python run_osintai.py --retry-model RUN_ID --retry-limit 20 --retry-timeout 60
python run_osintai.py --analyze-only RUN_ID --evaluate
```

Retries make no page-fetch requests and issue at most one model request per selected page.
They preserve original model outputs and publish a new `model_retry_*/` directory. The
`model_retry_latest.json` pointer selects the overlay used by later offline analyses.
`--model`, `--prompt-profile`, and `--analysis-max-chars` also apply to retries.

Each analysis first writes a hidden `.analysis_*.incomplete/` staging directory. After
JSON validation and file flushing, it publishes an immutable `analysis_*/` bundle containing
the report, result files, summary, optional training export, and manifest. Consult
`analysis_latest.json` to locate the latest completed bundle; older root-level analysis
files belong to earlier versions and are not updated. Separate `analysis_*.status.json`
files record running, failed, or completed attempts. A hard process kill can leave a
running status and an incomplete directory; these are never selected by the latest pointer.
Completed bundles can report partial analytical coverage: completion means publication
succeeded. Manifests include full source-file hashes and configured limits, and a source
change during analysis prevents publication. POSIX directory metadata is fsynced; Windows
uses atomic replacements and flushed files without POSIX directory-fsync semantics.

Extraction logs each page before scanning and each analysis stage before starting.
`--analysis-max-chars 200000` controls the per-page text limit (default: 200,000 characters).
The summary records truncation and the number of pages beyond the 2,000-page extended-scan
limit; findings from truncated inputs represent partial coverage. Original saved text is
retained. Ctrl-C exits with status 130 and a recovery hint instead of a traceback.
Unicode domain extraction consumes candidate tokens in linear time to avoid regex hangs.

After the crawl completes, OSINTai runs a deterministic analysis stage over what the crawl
collected. It is fast, works fully offline, and is on by default; `--no-analysis` skips it.

Expand Down Expand Up @@ -368,6 +442,10 @@ Each crawl generates a timestamped directory under `data/runs/` with comprehensi
- **`graph_edges.jsonl`** - Graph relationships and connections

### Analysis Results (unless `--no-analysis`)

These files live in the completed `analysis_*/` bundle selected by
`analysis_latest.json`, or under `reanalysis_*/analysis_*/` for offline recovery.

- **`analysis_report.txt`** - Findings separated by origin, correlations, timeline, hypotheses, leads
- **`findings.jsonl`** - Every finding with priority, evidence, sources, confidence, and next step
- **`correlations.jsonl`** - Scored candidate entity links with the evidence URLs behind each
Expand Down Expand Up @@ -705,7 +783,25 @@ pip install black flake8 pytest mypy

## Changelog

### v4.0.0 (2026-08-14) - Current Release
### v4.2.0 (2026-09-10) - Current Release
- Process deadlines for extraction and expensive analysis stages, with child cleanup
- Secret-free content-addressed extraction checkpoints and explicit coverage statistics
- Linear token scans for email, domain, credential, and JWT candidates
- Per-model response quality counts and bounded saved-page retries
- Unicode-correct hunt offsets and URL provenance
- Budgeted correlation, set-backed source tracking, and bounded text caching
- Validated immutable report bundles, source hashes, and durable attempt status
- Offline acceptance tests, saved-crawl benchmarks, and a three-platform CI matrix

### v4.1.0 (2026-09-10)
- Fix catastrophic backtracking in Unicode domain extraction using a linear token scan
- Add offline `--analyze-only RUN_ID` recovery into a fresh results directory
- Add per-page progress, bounded text reads, coverage statistics, and extraction failure isolation
- Preserve HTML element boundaries to prevent concatenated URL/label artifacts
- Handle interruption cleanly; write complete summary counts and empty result files
- Respect explicit `--flag=value` profile overrides and reject invalid numeric limits

### v4.0.0 (2026-08-14)
- **Evidence-Labelled Analysis**: Findings distinguish observed, derived, model-assisted, and hypothetical statements
- **Expanded Deterministic Checks**: Homoglyphs, sensitive infrastructure, secret presence, generated text, temporal gaps, and outliers
- **Cross-Source Intelligence**: Entity normalization, evidence-backed candidate correlations, timelines, hypotheses, and pivot leads
Expand Down Expand Up @@ -755,4 +851,4 @@ Built for the OSINT community with contributions from security researchers, digi

---

*OSINTai v4 - Illuminating the shadows of open source intelligence.*
*OSINTai v4.2.0 - Illuminating the shadows of open source intelligence.*
2 changes: 1 addition & 1 deletion src/osintai/__init__.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
__version__ = "4.0.0"
__version__ = "4.2.0"

__all__ = [
"cli",
Expand Down
Loading
Loading