fix(opt): preserve attention_mask=None in create_causal_mask for SDPA torch.compile - #48939
Open
ArjunPakhan wants to merge 8 commits into
Open
ArjunPakhan wants to merge 8 commits into
ArjunPakhan wants to merge 8 commits into
Conversation
…_initialized check
…f _is_hf_initialized is set
…skipping _initialize_weights
…nitialize_weights
Contributor
|
[For maintainers] Suggested jobs to run (before merge) run-slow: qwen3_omni_moe |
Contributor
CI recapDashboard: View test results in Grafana |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #48924
Summary
When
attention_maskisNone, OPT was instantiating an all-onesattention_masktensor for positional embeddings and overwriting the variable before callingcreate_causal_mask. This causedcreate_causal_maskto receive a concrete 2D tensor mask instead ofNone, preventing SDPA from relying onis_causal=Trueand triggering Triton kernel fallbacks undertorch.compile.Solution
Decoupled positional embedding mask calculation into
pos_attention_mask. This ensuresattention_maskremainsNonewhen passed intocreate_causal_mask, preserving the fast-path SDPA tracing logic.Testing
Ran local OPT unit test suite:
pytest tests/models/opt/test_modeling_opt.py-> 139 passed.