Walkyrie‑1.3B‑v1.0 (Preview) – Text‑to‑Image Diffusion Model
Test Walkyrie‑1.3B‑v1.0 to gauge image quality and anatomy accuracy.
Run sample prompts and report any anatomical errors to the repo maintainers.
Summary
The author releases Walkyrie‑1.3B‑v1.0, a text‑to‑image diffusion model derived from Wan‑AI/Wan2.1‑T2V‑1.3B. The model’s UMT5 text encoder was pruned to approximately 1 B parameters, and the architecture was retrained for image generation, converting the original text‑to‑video pipeline into a high‑quality text‑to‑image one. This is an early release with only about 20 % of the planned training budget, intended for testing and community feedback. The author notes that anatomy remains a significant issue, a common problem with small‑scale models. The post encourages community support to help improve the model.
Walkyrie‑1.3B‑v1.0 demonstrates the feasibility of repurposing a text‑to‑video model for image generation, but the limited training budget results in noticeable quality gaps, especially in anatomy. The author’s call for feedback highlights the collaborative nature of open‑source model development.
Developers should run sample prompts to assess image quality and anatomy accuracy, and report any issues back to the repository maintainers.
Key changes
- Derived from Wan2.1‑T2V‑1.3B.
- Text encoder UMT5 pruned to ~1 B parameters.
- Re‑trained for image generation.
- Early release with ~20 % training budget.
- Expected quality improvement with more training.
- Current anatomy issues noted.
- Community feedback encouraged.