MemoBench is a benchmark for evaluating visual memory in world generation models. It tests whether models can faithfully regenerate objects after they leave and re-enter the camera's field of view, following a V-D-R (Visible → Disappeared → Reappear) paradigm.
The benchmark includes 360 clips (196 synthetic + 164 real-world) spanning diverse scenes and physical state-change processes, evaluated across 14 metrics covering visual quality, temporal consistency, geometric fidelity, object permanence, camera controllability, and VQA-based reasoning.
- Installation
- Data Preparation
- Inference Protocol
- Output Directory Structure
- Running Evaluation
- Leaderboard
- Metrics Reference
- License
- Citation
- Linux (tested on Ubuntu 22.04)
- CUDA 12.x compatible GPU(s)
- Conda or Miniconda
conda create -n memobench python=3.11 -y
conda activate memobench
# PyTorch (adjust cu128 to match your CUDA driver version)
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
# Dependencies
pip install -r evaluation/requirements.txt
# MapAnything (required for Camera Controllability — estimates camera poses from generated videos)
git clone https://github.com/facebookresearch/map-anything.git third_party/map-anything
cd third_party/map-anything
pip install -e .
cd ../..
# SAM-3 (required for Object Revisit Score)
git clone https://github.com/facebookresearch/sam3.git third_party/sam3
cd third_party/sam3
pip install -e .
cd ../..conda activate memobench
python -c "import torch; import open_clip; import lpips; print('OK')"On recent environments (Python ≥ 3.12, setuptools ≥ 70, transformers ≥ 4.5x) the ImageReward metric can fail to load due to two upstream packages that haven't kept up. Both are one-line fixes inside your site-packages:
clippackage —ImportError: cannot import name 'packaging' from 'pkg_resources':
SITE=$(python -c "import clip, os; print(os.path.dirname(clip.__file__))")
sed -i 's/from pkg_resources import packaging/import packaging.version\nimport packaging/' $SITE/clip.py- ImageReward's bundled BLIP —
ImportError: cannot import name 'apply_chunking_to_forward' from 'transformers.modeling_utils'(these utils moved totransformers.pytorch_utils):
SITE=$(python -c "import ImageReward, os; print(os.path.dirname(ImageReward.__file__))")
python - "$SITE/models/BLIP/med.py" <<'EOF'
import sys
f = sys.argv[1]; s = open(f).read()
s = s.replace(
"from transformers.modeling_utils import (\n PreTrainedModel,\n apply_chunking_to_forward,\n find_pruneable_heads_and_indices,\n prune_linear_layer,\n)",
"from transformers.modeling_utils import PreTrainedModel\nfrom transformers.pytorch_utils import (\n apply_chunking_to_forward,\n find_pruneable_heads_and_indices,\n prune_linear_layer,\n)")
open(f, "w").write(s)
print("patched", f)
EOFIf ImageReward is unavailable, run_eval.py still runs — the ImageRewardScore
column is simply left empty (a warning is printed).
Download the MemoBench dataset and place it under the data/ directory:
MemoBench/
├── data/
│ ├── Synthetic_processed/ # Synthetic GT data (196 clips)
│ │ ├── Barnyard/
│ │ │ ├── Barnyard_001/ # (digit suffix below = clip number, e.g. RGB5/ for _005)
│ │ │ │ ├── RGB1/ # GT RGB frames; first frame = reference start frame
│ │ │ │ ├── Depth1/ # GT metric depth maps
│ │ │ │ ├── 1/
│ │ │ │ │ ├── intrinsics.npy # (N_gt, 4) [fx, fy, cx, cy]
│ │ │ │ │ └── poses.npy # (N_gt, 4, 4) cam extrinsics (c2w)
│ │ │ │ └── camera_full1.csv # raw UE5 camera track
│ │ │ ├── Barnyard_002/
│ │ │ └── ...
│ │ ├── CityPark/
│ │ ├── CityStreet/
│ │ ├── Japanese_Street/
│ │ ├── Miami/
│ │ ├── Nordic/
│ │ ├── NYstreet/
│ │ ├── OldFactory/
│ │ ├── WinterTown/
│ │ ├── ZenGarden/
│ │ └── Synthetic_ExitReenter.xlsx # V-D-R phase boundaries (synthetic)
│ │
│ ├── Real_Raw/ # Real-world GT data (164 clips)
│ │ ├── 001/
│ │ │ ├── 001.mp4 # GT video (also defines the GT frame count)
│ │ │ ├── FRAMES/ # extracted GT frames (frame_0001.jpg, ...)
│ │ │ ├── depth/ # per-frame depth maps
│ │ │ ├── raw/
│ │ │ │ └── 001_mapanything.npz # MapAnything GT camera: cam_c2w + intrinsic
│ │ │ └── thresholded/
│ │ │ └── 001_mapanything.npz # alternative estimate (not used by eval)
│ │ ├── 002/
│ │ ├── ...
│ │ └── Real_ExitReenter.xlsx # V-D-R phase boundaries (real)
│ │
│ ├── sam3_metadata/ # Per-clip prompts, first-frame paths, ORS subjects
│ │ ├── synthetic_metadata.csv # scene, video_id, subject, exr_path (first-frame path), prompt
│ │ └── real_metadata.csv # video_id, subject, full_image_path (first-frame path), prompt
│ │
│ └── vqa_questions/ # Pre-generated & filtered VQA
│ ├── Barnyard-1-questions.csv # Per-clip question sets (45 files)
│ ├── Barnyard-2-questions.csv
│ ├── ...
│ └── failure-cases.csv # VQA clip list: scene, video_id, folder, hint, prompt
Key files explained:
Synthetic_ExitReenter.xlsx(underSynthetic_processed/) /Real_ExitReenter.xlsx(underReal_Raw/): Define the V-D-R phase boundaries per clip — the frame where the object leaves the FOV and the frame where it returns. The two files use different column names: the synthetic sheet usesexits/re-enter, the real sheet usesexit/reenter. These are in GT frame space and are automatically scaled to match the generated video's frame count. These sheets are used at evaluation time only — they are not model inputs. Note thatReal_ExitReenter.xlsxhas nopromptcolumn: the per-clip generation inputs for the Real split (text prompt + reference first frame) live indata/sam3_metadata/real_metadata.csv(see Inference Protocol).{n}/poses.npy(synthetic) /Real_Raw/{id}/raw/{id}_mapanything.npz(real): GT camera poses. Synthetic poses come from the UE5 renderer (with the raw camera track incamera_full{n}.csv). Real-world GT poses and intrinsics are pre-estimated using MapAnything and shipped per clip in theraw/subfolder — you do not need to run MapAnything on GT videos yourself.sam3_metadata/*.csv: Per-clip metadata serving two roles. (1) Generation inputs: thepromptcolumn holds the full text prompt for each clip, and the reference first frame is given byexr_path(synthetic — historical column name, the path points to the firstRGB{n}/JPG) /full_image_path(real) — see Inference Protocol. (2) Evaluation: thesubjectcolumn (e.g., "fox", "person in silver suit") is used as the SAM-3 text prompt for ORS.vqa_questions/*-questions.csv: Pre-generated and filtered Yes/No question banks (6 per dimension × 4 dimensions).vqa_questions/failure-cases.csv: The clip list the VQA step iterates (columns:scene,video_id,hint,folder,Finished,prompt).folderis the generated-frames directory name for each clip;llm-vqa.pyreads this file by default via--cases-csv.
How to generate videos for MemoBench with your own model. The rule: every model is conditioned on the same first frame + text prompt; models that support camera conditioning additionally receive the GT trajectory. Keeping these inputs identical across models is what makes results comparable.
Per-clip generation inputs:
| Split | Text prompt | Reference first frame | Camera (optional) |
|---|---|---|---|
| Synthetic | prompt in data/sam3_metadata/synthetic_metadata.csv (same text as Synthetic_ExitReenter.xlsx) |
exr_path in the same CSV (points to the first RGB{n}/ JPG) |
intrinsics.npy + poses.npy ((N,4,4) c2w) per clip |
| Real | prompt in data/sam3_metadata/real_metadata.csv |
full_image_path in the same CSV |
Real_Raw/{id}/raw/{id}_mapanything.npz (cam_c2w + intrinsic) |
By model interface:
- CI2V (camera-controllable image-to-video): first frame + prompt + GT camera trajectory.
- Explicit-3D view synthesis: first frame + explicit GT camera poses.
- I2V / TI2V (no camera support): first frame + prompt only. Do not feed camera information. Such models are still scored on Camera Controllability and Instruction Following (and will score low there) — that is expected and reported as-is.
Notes:
- The V-D-R phase boundaries (
*_ExitReenter.xlsx) are evaluation-time only and are automatically rescaled to your generated video's frame count — you do not need to align frames to them during generation. - The GT camera poses serve double duty: a generation input for camera-conditioned models, and the reference for the Camera Controllability metric (poses are re-estimated from your generated video with MapAnything and compared against GT) for all models.
After generation, arrange your frames as described in Output Directory Structure and run the three-step evaluation below.
Place your model's generated video frames in the following structure:
output/{model_name}/
├── Synthetic/
│ ├── {Scene}/
│ │ ├── {Scene}_{NNN}/
│ │ │ └── frames/
│ │ │ ├── 00000.png
│ │ │ ├── 00001.png
│ │ │ └── ...
│ │ └── ...
│ └── ...
│
└── Real/
├── {NNN}/
│ └── frames/
│ ├── 00001.png
│ ├── 00002.png
│ └── ...
└── ...
Requirements:
- Format: PNG images, one frame per file.
- Naming: Zero-padded integers (e.g.,
00000.png). The evaluation pipeline sorts frames lexicographically. - Synthetic scenes: Barnyard, CityPark, CityStreet, Japanese_Street, Miami, Nordic, NYstreet, OldFactory, WinterTown, ZenGarden (and others). Clip IDs follow
{Scene}_{NNN}(e.g.,Barnyard_001). - Real clips: Numbered directories (e.g.,
001/,009/,120/).
The full evaluation pipeline has three steps:
Step 1: run_eval.py → automated metrics CSV
Step 2: compute_ors.py → ORS scores CSV
Step 3: llm-vqa.py → VQA scores CSVs
leaderboard.py → unified leaderboard
All commands below assume you are in the MemoBench/ root directory with the memobench environment activated.
conda activate memobenchComputes the automated metrics per clip: VisualQuality, MotionSmoothness, ObjIdentityConsistency, Geo3DConsistency, CameraControllability, ImageRewardScore, and pixel fidelity (PSNR, SSIM, LPIPS) — reported overall and per V/D/R phase.
Evaluate synthetic clips:
python evaluation/run_eval.py \
--mode synthetic \
--gen_root output/my-model/SyntheticEvaluate real-world clips:
python evaluation/run_eval.py \
--mode real \
--gen_root output/my-model/RealEvaluate both at once:
python evaluation/run_eval.py \
--mode both \
--gen_root_syn output/my-model/Synthetic \
--gen_root_real output/my-model/RealEvaluate a single clip (useful for debugging):
python evaluation/run_eval.py \
--mode synthetic \
--gen_root output/my-model/Synthetic \
--clip Barnyard_001For all command-line options, output format, and implementation details, see Step 1 Supplementary Details.
ORS measures whether the model regenerates the target object when the camera returns in the R-phase. It uses SAM-3 text-prompted segmentation with the object description from data/sam3_metadata/.
Before running, register your model in compute_ors.py by adding an entry to the MODEL_CONFIGS dict:
MODEL_CONFIGS = {
# ...existing models...
"MyModel": {
"synthetic": ["output/my-model/Synthetic"],
"real": ["output/my-model/Real"],
},
}Run ORS for your model:
python evaluation/compute_ors.py --model-name MyModelRun ORS for all registered models:
python evaluation/compute_ors.py --model-name allCustom output directory:
python evaluation/compute_ors.py \
--model-name MyModel \
--output-dir results/ors/Resume from partial run (skip already-scored clips):
python evaluation/compute_ors.py \
--model-name MyModel \
--skip-existingSAM-3 checkpoint location: By default, compute_ors.py looks for SAM-3 at third_party/sam3/. To override, set the SAM3_DIR environment variable:
SAM3_DIR=/path/to/sam3 python evaluation/compute_ors.py --model-name MyModelFor all command-line options, output format, and implementation details, see Step 2 Supplementary Details.
The VQA pipeline uses a Vision-Language Model (VLM) to answer pre-generated Yes/No questions about each generated video. It evaluates four dimensions: Instruction Following, Object & Background, Continuity of Memory, and Physics Adherence.
Prerequisites:
- A Gemini API key (set as environment variable
GEMINI_API_KEY) - Pre-generated question banks in
data/vqa_questions/(provided with the dataset; to regenerate, seeevaluation/vqa/llm-judger.py)
Before running, register your model's video directories in evaluation/vqa/llm-vqa.py by adding to MODEL_CONFIGS:
MODEL_CONFIGS = {
# ...existing models...
"MyModel": [
"output/my-model/Synthetic",
"output/my-model/Real",
],
}Run VQA evaluation:
export GEMINI_API_KEY="your-api-key-here"
python evaluation/vqa/llm-vqa.py \
--model-name MyModelResume from partial run:
python evaluation/vqa/llm-vqa.py \
--model-name MyModel \
--skip-existingRun all registered models:
bash evaluation/vqa/run_all_vqa.shFor all command-line options, output format, and implementation details, see Step 3 Supplementary Details.
After completing all three evaluation steps, aggregate results into a unified leaderboard:
python leaderboard/leaderboard.py \
--eval-dir outputs/ \
--ors-dir ors_results/ \
--vqa-dir vqa_results/ \
--output leaderboard.csvThe leaderboard script automatically:
- Loads the latest eval CSV per model from
--eval-dir/{model}/eval_*.csv - Loads ORS scores from
--ors-dir/{model}/ors_scores.csv - Parses per-clip VQA scores from
--vqa-dir/{model}/*.csv - Merges by model name (handles different naming conventions automatically)
- Converts all scores to their display scale (ORS: 0–1 → 0–100; VQA: 0–1 → 0–100)
Join the leaderboard:
🏆 Check out our Leaderboard on Huggingface for the latest rankings and numerical results!
How to join the leaderboard:
- Submit Results (Highly Recommended): Follow the evaluation instructions above to generate your
leaderboard.csvfile. Email it to tonychen54816@gmail.com, and the MemoBench Team will update the leaderboard. - Submit Video Samples: Alternatively, you can send us your raw video samples for us to evaluate. Please note that processing times for this method will vary depending on our team's bandwidth and resources.
For detailed descriptions of all 14 metrics (backbones, phases, score ranges), see the Metrics Reference.
This project is licensed under the MIT License. See LICENSE.md for details.
If you find MemoBench useful in your research, please cite:
@article{chen2026memobench,
title={MemoBench: Benchmarking World Modeling in Dynamically Changing Environments},
author={Chen, Haoyu and Zhou, Kaichen and Hua, Hang and Zhang, Kaile and Qian, Jingwen and Ma, Wufei and Chen, Haonan and Liu, Chunjiang and Zhao, Yizhou and Wang, Xiaoyuan and others},
journal={arXiv preprint arXiv:2606.27537},
year={2026}
}