Skip to content

protenix-v2 has no end-to-end pipeline #5

Description

@ssiddhantsharma

Summary

protenix-v2 is listed as a supported model, but there is no end-to-end pipeline for it: build_processor(model_source="protenix-v2") cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.

What I found

Using the same EngineProcessorConfigbuild_processor path that folds Boltz-2 end-to-end, Protenix-v2 does not run: ProtenixFactory's tokenizer / feature_factory / postprocessor raise NotImplementedError (the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.

If you instead try to drive the optimized Protenix model directly, the forward expects an input_feature_dict on a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys like d_lm / v_lm and drops profile / deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.

Reproduce

from bionemo_ir.registry import register_all_factories; register_all_factories()
from bionemo_ir.pipeline.processor.engine_proc import EngineProcessorConfig, build_processor

cfg = EngineProcessorConfig(
    model_source="protenix-v2",
    runtime_args={"diffusion_samples": 1, "num_sampling_steps": 200, "recycling_steps": 10},
    writer_stage={"output_path": "/out", "format": "cif"},
)
proc = build_processor(cfg)   # protenix-v2: tokenizer/feature_factory/postprocessor NotImplementedError

(The identical pattern with model_source="boltz-2" folds fine.)

Ask

  1. Is an end-to-end Protenix-v2 pipeline planned (tokenizer + featurizer + postprocessor wired into the stage framework, like Boltz-2)?
  2. If not, would a PR porting the OSS ByteDance Protenix featurization into the stage framework be welcome? Happy to contribute if the direction is wanted.
  3. Separately, could the expected input_feature_dict schema for the optimized Protenix forward (the d_lm/v_lm keys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?

Environment

bionemo-ir 0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions