Storifai
In Development
Three story versions from a single image upload.
Experimental — not for real-world use·Try the live demo →
Expected Q4 2026

The Problem
Image captioning models generate generic descriptions. They don't tell stories — and they don't give you creative control.
The Approach
Upload an image. Storifai uses CLIP + Cross-Image Attention + Transformer Decoder to generate three tonally distinct story versions. Built as a CS747 computer vision project with BLEU benchmark evaluation.
Ethics & Disclosure
- Evaluated against a baseline to measure actual improvement
- Transparent about model architecture and limitations
- Research-grade — not production storytelling software
Highlights
- Three story variants per upload (literary, journalistic, imaginative)
- CLIP + Cross-Image Attention + Transformer Decoder architecture
- BLEU score benchmarking against baseline
- CVPR-format research paper
Get notified when it ships
No spam. One email when it's ready.