Aethon AI Launches a Production AI-Visibility Platform in 12 Weeks

This case shows how a focused product-engineering team can take a new categoryfrom a whiteboard thesis to a live, multi-tenant SaaS platform, one that listensto the AI conversations buyers are actually having, and turns what it hears intoa ranked action plan.

5 AI assistants monitored on an automated daily cycle

~12 weeks from first commit to production platform

295 automated tests shipped alongside the product

Customer
Aethon AI
Website
https://aiaethon.com/
Platform
Next.js 16 · React 19 · PostgreSQL 18 · FastAPI · Temporal · LangChain
Industry
Marketing intelligence / AI SaaS

Daniel Arons
Co-founder - Aethon AI

Aethon AI is the platform behind Contextual AI Presence Mapping — a category the company defined to answer a question classic analytics cannot: when a buyer asks an AI assistant for a recommendation, does your brand make it into the answer?

When they came to us, Aethon had identified the gap and named the category. What they did not have was the platform underneath it.

Search is no longer the front door. Buyers open ChatGPT, Claude, Gemini, Perplexity or Grok, describe their situation in plain language, and take the recommendation they are given. There is no ad slot in that answer and no blue-link ranking to optimise. You are either named or you are not.

Key Challenges

The measurement target moves and disagrees with itself across five assistants

Buyers decide over multi-turn conversations, not single prompts

Setup friction would kill adoption before the first measurement

Measurement alone is not a product — a percentage tells a CMO nothing to do on Monday

Where They Were Before

Aethon AI had identified the gap and named the category. What they needed was theplatform underneath it, and the requirements were unusually demanding for anearly-stage build:

  • The measurement target moves and disagrees with itself. Five AI assistants, each with its own model, its own answer, and different results on different days. A single sample proves nothing.

  • A single prompt isn't how buyers behave. Real purchase decisions happen over a conversation, and a brand's position can shift dramatically between the first question and the final decision.

  • Setup friction would kill adoption. The obvious design — ask the marketer to write out the queries to track — puts a blank page in front of a new user in the first 60 seconds.

  • Measurement alone isn't a product. A dashboard reporting that you appear in 31% of answers tells a CMO nothing about what to do on Monday.

  • It had to be real SaaS from day one. Multi-tenant, auditable, and safe to put in front of a paying customer.
The Problem

Methodology

We treated the AI assistants as an unreliable, rate-limited, non-deterministic data source — because that is exactly what they are — and built the platform around that reality rather than against it. Every measurement is a durable, retryable, independently observable unit of work. Every AI-authored output is validated before it is allowed to become a fact in the product.

Strategic Rationale

In a monitoring product, the expensive failures are the silent ones. A missing row that looks like an absent brand, a prompt tweak nobody can trace, a score no one can interrogate — each erodes trust faster than any feature can rebuild it. The architecture therefore prioritised observability and reversibility over raw speed to first demo.

1. Onboarding as inference, not data entry

The user types one thing: their website URL. The platform reads the site, infers the brand's thesis, its buyer personas, the life moments those buyers are in, the conversations they are likely having with an assistant, and the real competitors — then presents it all for review. About twenty seconds. Confirming the review provisions the account, persists the moments, registers a daily schedule for every query, and fires the first sixty measurement calls. The draft is stored server-side, so a user can close the tab mid-review and pick up exactly where they left off.

2. Durable orchestration instead of a job queue

Every query-by-platform measurement is its own Temporal workflow with a retry policy. Sweeps run every 24 hours; industry benchmarks every 30 days. When a provider fails, the run is recorded as failed with the provider's actual error message — an operator can tell the difference between a brand that was not mentioned and a model id that changed overnight.

3. Multi-turn conversation capture

One click runs a buyer-shaped dialogue across all five assistants in parallel. Each turn is captured in full: the answer, whether the brand appeared, where it ranked, and every competitor named. What this exposes is recommendation drift — the way a brand's standing moves across a conversation, which a single-prompt rank check cannot see at all.

4. Prompt engineering as versioned infrastructure

The prompts driving every AI feature are not constants in the source code. They live in the database as versioned skills: published versions are append-only, services compose skills through named slots, weighted A/B tests can run between variants with a written hypothesis and a success metric, and every execution records which composition served it, whether validation passed, and how long it took. Prompts can be tuned, split-tested and rolled back without shipping code.

5. A strategy engine that shows its work

Measurement feeds a pipeline that scores each gap — a moment where the brand is absent, weak, or slipping on a given platform — across five dimensions, diagnoses the likely cause, clusters the gaps into projects, and sequences them into three phases from fast owned-channel wins through to long-horizon authority building. Each score carries its own breakdown, so the product can explain in plain language why a gap ranked where it did.

6. Two products in one surface

The daily read is a magazine — moment of the week, this week in numbers, where you are losing, where you could win next — designed to be read like a newspaper by an executive. Behind it sits a full analytics dashboard for the operator who needs to drill into a single moment, provider by provider, turn by turn. Same data, two completely different reading postures.

Key Strategies

Onboarding as inference — one URL to a provisioned, monitoring account

Durable orchestration with per-provider failure visibility

Multi-turn conversation capture to expose recommendation drift

Versioned prompt infrastructure and an interrogable strategy engine

Our Approach

The build moved through four distinct phases:

  • April 2026 — Architecture, multi-tenant data model, first provider adapters, first monitoring loop

  • May 2026 — Temporal orchestration and scheduling, landing-page generation with a separate QA validation contract, the database-backed skills runtime

  • June 2026 — Strategy engine, multi-turn conversation capture, competitor tracking

  • July 2026 — v2 platform: inference-driven onboarding, the For You brief, moment deep-dives with per-provider tabs

Each phase built on infrastructure decisions made in the one before it, rather than retrofitting them later.

Evolution Over Time

Deliverables:

  • Multi-tenant SaaS platform architecture and data model
  • 28-table schema with soft deletes throughout
  • Five live LLM provider integrations, nine modelled in the data layer
  • Temporal-based durable orchestration and scheduling
  • Database-backed versioned prompt runtime with A/B testing
  • Gap scoring and strategy sequencing engine
  • Executive brief surface and full analytics dashboard
  • 295 automated tests

Channels / platforms:

  • Next.js 16 / React 19 front end
  • Python / FastAPI orchestration back end
  • PostgreSQL 18 data layer
  • Temporal, LangChain
Execution

The Step Labs Effect

This case shows how a focused product-engineering team can take a new categoryfrom a whiteboard thesis to a live, multi-tenant SaaS platform, one that listensto the AI conversations buyers are actually having, and turns what it hears intoa ranked action plan.

Key Outcomes

5 AI assistants monitored on an automated daily cycle

~12 weeks from first commit to production platform

295 automated tests shipped alongside the product

Data-Led Decisions
Built to Grow
World-Class Delivery
Rapid Optimization
Trusted by Leaders
Engineered Growth
Data-Led Decisions
Built to Grow
World-Class Delivery
Rapid Optimization
Trusted by Leaders
Engineered Growth