Summary
protenix-v2 is listed as a supported model, but there is no end-to-end pipeline for it: build_processor(model_source="protenix-v2") cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.
What I found
Using the same EngineProcessorConfig → build_processor path that folds Boltz-2 end-to-end, Protenix-v2 does not run: ProtenixFactory's tokenizer / feature_factory / postprocessor raise NotImplementedError (the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.
If you instead try to drive the optimized Protenix model directly, the forward expects an input_feature_dict on a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys like d_lm / v_lm and drops profile / deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.
Reproduce
from bionemo_ir.registry import register_all_factories; register_all_factories()
from bionemo_ir.pipeline.processor.engine_proc import EngineProcessorConfig, build_processor
cfg = EngineProcessorConfig(
model_source="protenix-v2",
runtime_args={"diffusion_samples": 1, "num_sampling_steps": 200, "recycling_steps": 10},
writer_stage={"output_path": "/out", "format": "cif"},
)
proc = build_processor(cfg) # protenix-v2: tokenizer/feature_factory/postprocessor NotImplementedError
(The identical pattern with model_source="boltz-2" folds fine.)
Ask
- Is an end-to-end Protenix-v2 pipeline planned (tokenizer + featurizer + postprocessor wired into the stage framework, like Boltz-2)?
- If not, would a PR porting the OSS ByteDance Protenix featurization into the stage framework be welcome? Happy to contribute if the direction is wanted.
- Separately, could the expected
input_feature_dict schema for the optimized Protenix forward (the d_lm/v_lm keys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?
Environment
bionemo-ir 0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.
Summary
protenix-v2is listed as a supported model, but there is no end-to-end pipeline for it:build_processor(model_source="protenix-v2")cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.What I found
Using the same
EngineProcessorConfig→build_processorpath that folds Boltz-2 end-to-end, Protenix-v2 does not run:ProtenixFactory's tokenizer / feature_factory / postprocessor raiseNotImplementedError(the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.If you instead try to drive the optimized Protenix model directly, the forward expects an
input_feature_dicton a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys liked_lm/v_lmand dropsprofile/deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.Reproduce
(The identical pattern with
model_source="boltz-2"folds fine.)Ask
input_feature_dictschema for the optimized Protenix forward (thed_lm/v_lmkeys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?Environment
bionemo-ir0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.