2BFT Studio

The visual lab.

We test AI models so you know which one to use. Public results, no paywall, no hype. If a tool doesn't work, we say so. If it works, we tell you how we use it.

What we test

The 8 tools in the lab.

Nano Banana 2 + Lite

Nano Banana 2 (released Feb 26, 2026, 4K in 4–6s) is the default in Gemini, Search AI Mode, Lens, Vertex AI, AI Studio, Antigravity and Flow. Nano Banana 2 Lite (released June 30, 2026) is even faster and cheaper. Our special weapon — used to render Stashed before the first sample was stitched.

Veo 3.1 + Lite

Veo 3.1 is the best AI video after OpenAI retired Sora 2 consumer — 1080p, native audio, strongest physics consistency. Veo 3.1 Lite is the public preview on Vertex AI, most cost-efficient. 60-second storyboards for brands like Maruthi Jewellers and SN Bags.

Higgsfield

Character animation. When you need an avatar to do something specific — not a stock image, an actual moment.

GPT-5.6 (new)

OpenAI's GPT-5.6, released July 9, 2026 — three tiers (Sol, Terra, Luna), strong SWE-bench Pro. Now default in ChatGPT. The first thing we test on day one of every release.

Claude Sonnet 5 (new)

Anthropic's Claude Sonnet 5, released June 30, 2026. 1M context window, 128K max output. Best in class for writing style and instruction-following. Our default for copy and reasoning.

Gemini 3.1 / 3.5 Pro (new)

Google's Gemini 3.1 Pro leads on hardest reasoning (94.1% HumanEval, 2M+ context). Gemini 3.5 Pro was cleared for July GA after limited preview. We test both against the open-weights competition.

OpenClaw (new)

Self-hosted AI agent framework — 210K+ GitHub stars (the highest of any project in 2026), no external API calls, free and open-source. The first thing we run before any other agent stack.

ElevenLabs v3

Voice AI. ElevenLabs Voice Engine v3 — improved prosody, multilingual. Audio content at scale: voiceovers, B-roll narration, podcast intros.

Recent experiments

What's on the bench right now.

GPT-5.6 vs Claude Sonnet 5 vs Gemini 3.1 Pro — same 5 prompts

Released within 9 days of each other in late June / early July 2026. We ran the same 5 prompts across all three on day one. Results + winner for Indian-context use cases.

PUBLISHED →

Nano Banana 2 vs. Midjourney v6 — product photography

Same prompt, same lighting, same bag. The result: Nano Banana 2 wins for Indian context, Midjourney wins for global.

PUBLISHED →

Veo 3.1 vs. Sora — 10-second product clip

We rendered 50 clips across both. Here's what we use each for.

PUBLISHED →

OpenClaw in production — 3 months in

Self-hosted agent stack running 50+ integrations with zero external API calls. What broke. What worked. What we'd build differently.

PUBLISHED →

The Indian Model Lab — first benchmarks

Comparing Indian-focused AI models on real tasks. Public data, no paywall, no sponsor bias.

PUBLISHED →
Phase 4 of the Academy

The Indian Model Lab.

Compare Indian-focused AI models, contribute test cases, and help us build the benchmark. Results published publicly. No paywall.