feat: add PrivateUse1 backend extension support - #206
Conversation
| auto hook = std::make_unique<infini_train::autograd::AllReducePostAccumulateHook>( | ||
| function::ReduceOpType::kAvg, ddp_pg_); | ||
| const auto reduce_op | ||
| = ddp_config.average_in_collective ? function::ReduceOpType::kAvg : function::ReduceOpType::kSum; |
There was a problem hiding this comment.
关于 reduce_op 类型和架构后端实现不是强相关的,是不是单独提一个PR
| } else { | ||
| bucket.work = ddp_pg->AllReduce(bucket.contents, function::ReduceOpType::kAvg, true); | ||
| const auto reduce_op | ||
| = ddp_config_.average_in_collective ? function::ReduceOpType::kAvg : function::ReduceOpType::kSum; |
| @@ -143,18 +145,11 @@ void DeviceGuardImplRegistry::Register(Device::DeviceType type, std::unique_ptr< | |||
| LOG(FATAL) << std::format("DeviceGuardImpl for type {} already registrered", static_cast<int>(type)); | |||
| } | |||
|
|
|||
There was a problem hiding this comment.
删除单加速器后端限制后,Tensor::To(Device) 中原有的跨后端复制路径就可以覆盖到了, tensor.cc::161 存在一个问题,第二步 H2D 复制根据 buffer_device获取impl,本来应该使用目标 device来获取tmpl。这里comment作记录,可以另外PR修复,加单元测例覆盖一下。
There was a problem hiding this comment.
InfiniTrain 现在只允许 CPU + 一个 PrivateUse1 provider,应该还不涉及跨后端,但这里确实有潜在问题,留 FIXME 记录
| # ------------------------------------------------------------------------------ | ||
|
|
||
| add_library(infini_train STATIC ${SRC}) | ||
| add_library(InfiniTrain::infini_train ALIAS infini_train) |
There was a problem hiding this comment.
新增的 InfiniTrain::infini_train alias 是不是给外部 provider 直接链接使用的?目前 runtime、CCL 和 kernel 都依赖静态注册,而保证这些注册代码不被链接器裁掉的 --whole-archive 只加在 link_infini_train_exe() 里。如果外部工程直接 target_link_libraries(... InfiniTrain::infini_train),没有调用 link_infini_train_exe(),运行时报 runtime 或 kernel 未注册
There was a problem hiding this comment.
不是给外部 provider 直接链接使用的,provider 必须用 RegisterBackend(),我加一下约束注释。
| void RegisterFakeRuntime() { | ||
| CHECK_EQ(core::GetPrivateUse1BackendName(), "fake"); | ||
| CHECK_EQ(Device(Device::DeviceType::kPrivateUse1, 0).ToString(), "Device(fake, 0)"); | ||
| INFINI_TRAIN_REGISTER_DEVICE_GUARD_IMPL(Device::DeviceType::kPrivateUse1, FakePrivateUse1GuardImpl) |
There was a problem hiding this comment.
warning: unused variable ‘__infini_train_device_guard_registered__COUNTER__’ [-Wunused-variable]
236 | static const bool __infini_train_device_guard_registered##__COUNTER__ = []() { \
| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
这里会有一个告警,因为之前加了-Wunused 编译选项,是不是在 Register 宏里加一下[[maybe_unused]]。
同时也发现了一个问题,在宏里 __COUNTER__直接接触 ##,没有展开成数字,在此记录一下,后续另提PR修改。
| @@ -143,18 +145,11 @@ void DeviceGuardImplRegistry::Register(Device::DeviceType type, std::unique_ptr< | |||
| LOG(FATAL) << std::format("DeviceGuardImpl for type {} already registrered", static_cast<int>(type)); | |||
0c5953b to
4b1042c
Compare
- add a provider-neutral PrivateUse1 device type and registration API - validate runtime, kernel, and optional CCL backend registrations - initialize external device runtimes lazily on first use - support provider names in device parsing and display - require explicit autocast dtype for PrivateUse1 devices - allow examples to register an external backend before flag parsing - honor average_in_collective consistently across DDP paths - expose embeddable CMake targets and add fake backend tests
- separate test suite declaration from CPU, CUDA, and provider instantiation - support provider-injected registration, linkage, target names, and CTest labels - add provider-defined default autocast dtype and backend-neutral test helpers - generalize accelerator copy tests and remove the ineffective CUDA optimizer test - fix Exp/Add backward, CUDA bias reduction, and DDP device validation - scope runtime workarounds to the registered MACA backend - document external backend test integration and usage
4b1042c to
66cc91e
Compare
66cc91e to
54dc847
Compare
背景
InfiniTrain 原有设备体系只包含 CPU 和 CUDA。接入新后端时,需要在核心框架中增加厂商专属的设备枚举、runtime、CCL、kernel、测试和模型入口判断,导致核心代码与具体厂商耦合。
本 PR 引入通用的
DeviceType::kPrivateUse1扩展槽位。外部 Provider 可以注册自己的运行时和算子实现,MACA 后端则通过独立仓库接入,公共 CMake 中不包含 MACA SDK 或厂商构建选项。主要改动
PrivateUse1 注册接口
新增
PrivateUse1BackendRegistration和RegisterPrivateUse1Backend(),统一完成:macaDeviceGuardImplCclImplProvider 名称仅允许小写 ASCII 字母、数字和下划线,且不能占用
cpu、cuda。继续复用
REGISTER_KERNEL、INFINI_TRAIN_REGISTER_DEVICE_GUARD_IMPL和INFINI_TRAIN_REGISTER_CCL_IMPL。注册宏同时修正了唯一符号生成和未使用变量告警。PrivateUse1 Provider 至少需要提供:
CastFillNoOpForwardNoOpBackwardProvider 必须暴露可显式调用且幂等的注册入口,不能只依赖静态库中的文件级初始化。
DeviceGuardImpl::Initialize()改为首次使用时调用,允许 Provider 延迟初始化硬件 runtime。设备解析与模型入口
新增统一的
Device::ParseType():cpu映射到kCPUcuda映射到kCUDAprivateuse1映射到kPrivateUse1kPrivateUse1Device::ToString()使用注册后的厂商名称。AutocastGuard根据 Provider 注册信息取得 PrivateUse1 的默认计算类型,CPU 和 CUDA 的现有行为保持不变。GPT-2、Llama 3 和 Mixtral 使用
Device::ParseType()解析设备。外部工程可通过编译定义注入 Provider 头文件和注册入口,并在 gflags 校验--device前完成注册。并行模型使用用户选择的 accelerator backend,不再固定为 CUDA。运行脚本新增
DEVICE_BACKEND,默认值仍为cuda。GPT-2 和 Llama 3 暂时保留少量仅在 Provider 名称为
maca时启用的同步和进程退出 workaround,并保留 FIXME;这些逻辑不会影响其他 PrivateUse1 Provider。构建与静态注册
新增并明确以下 CMake 接口:
InfiniTrain::infini_train:供库和 Provider 使用的核心接口InfiniTrain::cpu_kernels:CPU kernel targetInfiniTrain::infini_train_executable:最终可执行文件的完整链接接口最终可执行文件通过 archive group 和
--whole-archive保留 DeviceGuard、CCL 和 kernel 的静态注册对象。Provider 可以通过自己的 executable interface 或EXTRA_ARCHIVES加入 Provider 注册 archive。InfiniTrain 作为 submodule 使用时不再构建自身 examples 和 tools,并隔离 glog 的测试选项,避免污染上层工程。
测试复用
公共测试的 suite 声明与设备实例化分离,每个测试二进制只实例化一个设备。CMake 为目标注入设备类型、GTest 前缀和
DEVICE_INDEX,其中设备序号默认是0。新增
infini_train_add_privateuse1_test_suites(),允许外部 Provider 复用全量公共测试:USE_CUDA=OFF,冲突时直接报错test_*_<BACKEND_NAME>Provider 测试test_main在测试环境初始化前显式调用ONLY_CUDA,copy 等测试改为通用 accelerator 测试Provider 测试使用厂商名作为 CTest label,例如 MACA 使用
ctest -L maca;不提供ctest -L privateuse1标签。新增无需真实硬件的 fake PrivateUse1 测试,覆盖注册校验、名称解析、默认 autocast dtype、延迟 runtime 初始化和基础 kernel 调度。
Test