A recent technical post demonstrates an end-to-end pipeline that integrates ERNIE-Image for text-to-image generation with InfiniteTalk for creating infinite-length videos, all deployed on domestic DCU (Deep Computing Unit) hardware. The pipeline covers the entire workflow from image generation to audio synchronization and digital human video creation, with complete code provided. This is significant because it showcases the maturation of domestic AI hardware ecosystems, enabling complex multimodal tasks that were previously reliant on foreign GPUs. The integration of these models on DCU suggests that Chinese AI hardware is becoming a viable platform for cutting-edge research and commercial applications in video generation. While the code is provided, the real signal is the feasibility of running such pipelines on domestic infrastructure, which has implications for AI sovereignty and cost reduction.
This article presents a full pipeline combining ERNIE-Image for text-to-image generation and InfiniteTalk for extending to infinite-length video, deployed on domestic DCU hardware. It includes complete code for generating synchronized audio and digital human videos. This signals a growing capability in domestic AI hardware ecosystems for complex multimodal tasks.