← back to posts

$ posts/football-content-agent.md

---
date:  2026-09-28
tag:   Showcase
read:  4 min
---

I built a multi-agent AI pipeline that runs an Instagram football page without me

Every morning at 7 AM ET, a five-stage multi-agent pipeline wakes up on Google Cloud, reads about 165 football stories from four sources, argues itself down to the best five, generates a graphic and caption for each, and posts them to Instagram. No human in the loop. The account is live at @football.news_updates_.

The idea started as one line in my projects list: an agent for football news (game updates, press conferences, team updates, no gossip) with data-rich stats and AI-generated visuals. It became my testbed for learning agent orchestration on something real: an audience, a daily deadline, and consequences when a stage fails.

What it does

Cloud Scheduler fires an HTTP POST at a Vertex AI Agent Engine deployment (built with Google’s Agent Development Kit). Inside, a SequentialAgent (ADK’s orchestrator that runs stages strictly in order) drives five stages, passing everything through session state, a shared key-value scratchpad each stage reads from and writes to:

fetcher          → ~165 raw ideas from football-data.org, NewsAPI,
                   Reddit r/soccer, and 5 RSS feeds (BBC, Sky, Guardian…)
idea_judge       → Gemini 2.5 Pro filters to ~25 visually interesting,
                   data-rich candidates
idea_ranker      → ranks and keeps the top 5
content_generator→ per idea: research (if needed) → plan the layout
                   → gpt-image-2 background → PIL text overlay → caption
publisher        → PNG→JPEG, crop to 4:5, upload to GCS,
                   publish via the Instagram Graph API

No agent knows about the others. Each stage reads named keys from session state and writes new ones: raw_ideas_json, then candidate_ideas_json, approved_ideas, final_posts, published_posts. That decoupling is what made the pipeline debuggable. I can rerun any stage against yesterday’s state.

The design patterns that earned their keep

Best model first, cheap models where proven. I followed the standard playbook: build with strong models everywhere (Gemini 2.5 Pro for judgment calls), establish baseline quality, then downgrade the easy seats. The researcher and caption writer run on 2.5 Flash because those tasks survived the downgrade; the judge, the ranker, and the content planner didn’t, so Pro keeps all three judgment seats.

Conditional agent activation. The researcher sub-agent (Flash plus Google Search grounding) only fires when an article’s text is under 800 characters or the judge flagged missing data. Most stories skip it. That’s an LLM call that exists only when it buys something.

A judge and a ranker, deliberately separate. “Pick the best 5 of 165” is two decisions wearing one trench coat: is this idea any good? and which good ideas win today? Splitting them let each prompt do one job, and the intermediate ~25-candidate list is where I look first when output quality dips. It’s the single-responsibility principle applied to prompts.

Fault isolation per source and per idea. One dead RSS feed doesn’t kill the fetch stage; one malformed idea doesn’t abort the other four posts. Every source and every idea gets its own try/except with a logged warning. When Reddit’s public API 403s (and it does), the pipeline shrugs.

Structured output with retry. The content planner must return a PostPlan Pydantic model. Gemini 2.5 Pro intermittently returns malformed JSON, so the generator retries up to 3 times with a 5-second backoff before dropping that idea. Retrying a schema failure is cheaper than a human fixing a broken post.

What broke (and what it taught me)

The commit log is honest about where the real work was: fixing horizontally-stretched images with a proportional cover-crop, bundling a Roboto font because the deployed container rendered tiny text, dropping a dead RSS feed, resolving dependency conflicts that only appeared on a fresh install. The agents were maybe a third of the effort. The rest was the unglamorous edge of shipping: image geometry, fonts, quotas, CI/CD.

Deployment runs through GitHub Actions with Workload Identity Federation (no service-account keys in the repo), and the pipeline health check I care most about is that a deploy updates the existing Agent Engine resource in place instead of silently creating a new one. A seen.json persisted across runs stops the pipeline from posting the same story twice. And the ops discovery that made the whole thing sustainable: an idle Agent Engine resource costs nothing, because it’s serverless, so “pausing” the product is just gcloud scheduler jobs pause football-content-agent-daily.

Would I build it as agents again?

Mostly yes. The multi-agent structure paid for itself in observability and fault isolation, not intelligence. The same pipeline as one giant prompt would be cheaper per run and impossible to debug. The honest caveat: a daily content pipeline has no latency requirement, which makes it a forgiving place to learn orchestration. The hard version of this problem is the same loop with a human waiting on the other end, which is what I took on next with aloud.

Repo: github.com/JigneshAmmineni/footballContentAgent-IG

EOF · back to posts