EchoFuture is a video masked autoencoder (VideoMAE) pre-trained on echocardiogram videos. EchoFuture-HFrEF is fine-tuned from EchoFuture to predict 10-year risk of heart failure with reduced ejection fraction (HFrEF) from a single echocardiogram clip.
This repository provides inference and evaluation code for both checkpoints, hosted on Hugging Face.
Python 3.10+ is required. A CUDA GPU is recommended for cohort-scale inference.
git clone https://github.com/VoyagerWSH/EchoFuture.git
cd EchoFuture
python -m venv .venv && source .venv/bin/activate
pip install -e .Model weights are downloaded from the Hub. If the repo is gated, export your token:
export HF_TOKEN=<your_token>Both checkpoints live in a single Hugging Face repository:
VoyagerWSH/EchoFuture/
├── pretrained/ # EchoFuture foundation model
│ ├── config.json
│ ├── model.safetensors
│ └── preprocessor_config.json
└── hfref/ # EchoFuture-HFrEF risk model
└── echofuture_hfref.pth
from transformers import VideoMAEForPreTraining
model = VideoMAEForPreTraining.from_pretrained(
"VoyagerWSH/EchoFuture",
subfolder="pretrained",
attn_implementation="sdpa",
)import torch
from huggingface_hub import hf_hub_download
from echofuture.model import EchoFutureHFrEF, strip_module_prefix
model = EchoFutureHFrEF(num_followups=10)
weights = hf_hub_download(
"VoyagerWSH/EchoFuture", filename="echofuture_hfref.pth", subfolder="hfref"
)
sd = torch.load(weights, map_location="cpu", weights_only=False)
if isinstance(sd, dict) and "state_dict" in sd:
sd = sd["state_dict"]
sd = strip_module_prefix(sd)
model.load_state_dict(sd, strict=True)
model.eval()The model expects input of shape (B, C=3, T=16, H=224, W=224) and outputs logits of shape (B, 10) — one value per yearly follow-up. Apply sigmoid to obtain probabilities.
Run EchoFuture-HFrEF on a cohort CSV:
echofuture-infer --data /path/to/cohort.csv --output-dir ./outputThis writes a timestamped predictions CSV plus fps.png and survival.png to the output directory.
The pipeline automatically applies inclusion criteria (baseline LVEF ≥ 50 %, A2C/A4C views, non-Doppler TTE, fps ≥ 15) and computes evaluation metrics including Uno's IPCW C-index, per-year AUROC/AUPRC with 95 % bootstrap CIs, and FPS-subgroup analysis. See docs/dataset.md for the required CSV schema.
Default settings are in echofuture/config_defaults/inference.yaml. Override selectively:
echofuture-infer --data cohort.csv --config my_config.yaml --output-dir ./outputKey options:
| Key | Default | Description |
|---|---|---|
model.repo_id |
VoyagerWSH/EchoFuture |
Hugging Face repo |
model.weights_file |
echofuture_hfref.pth |
Checkpoint filename |
train.batch_size |
64 |
Batch size for inference |
train.dtype |
float32 |
bfloat16 for faster GPU inference |
overall.seed |
12138 |
RNG seed for reproducibility |
echofuture/ # Python package
model.py # EchoFutureHFrEF architecture
dataset.py # Video dataset loader
inference.py # Batch inference pipeline
cli.py # CLI entry point
util.py # Data processing, metrics, video I/O
hf_hub_auth.py # Hub token resolution
config_defaults/ # Default YAML config
docs/
dataset.md # Input CSV column reference
assets/
EchoFuture.gif # Model overview animation
Under review.
