← Back to Work Jump to Entity Diagram ↓
Product Design · AI Tooling · UX Strategy

AutoQ — UX Strategy Document

Reserve the reset. Never wait on the wall again.

The wall
The "unlimited era" is over. Flat-rate AI plans have shifted to metered, rolling-window access — and users get cut off mid-thought, unable to see the meter or plan against it.
The reservation
AutoQ is a scheduling layer on the AI tools people already pay for. It reserves blocked work, fires it the instant the usage window resets, and bridges the wait on a free local model — turning dead time into draft time.
Year
2026 / Jul
Type
Product Design, AI Token Tool
Mechanic
Reset-timed execution + free local bridge
Status
UX Strategy

AutoQ sits on top of the AI tools people already pay for. When a user hits a usage limit, it decomposes the pending prompt into a task queue, schedules it to fire the moment the window resets, and — while the user waits — runs a draft pass on a free local web LLM (Qwen) so momentum never stops. You reserve AI work the way you'd reserve a table: set it, walk away, come back to a finished result.

From the wall to the reservation

The document moves from the market reality that creates the pain, to the users who feel it, to the flow and data model that resolve it, and finally to the metrics and strategy that make it a business.

01
Problem context & market research (2026)Why metered, windowed AI access is the new normal.
02
Target userNorth-American solo builders and async knowledge workers.
03
User flow — the reservation, step by stepThe happy path, with six strength moments today's tools can't match.
04
Entity diagram (ERD)Reservation → Task → (BridgeRun · Execution) is the spine.
05
Stakeholder mapThe system's actors, in primary / secondary / tertiary rings.
06
KPIsNorth Star & the funnel behind it.
07
As-Is → To-BeDead end becomes reservation.
08
Strategic perspectiveWhy it wins, how it's positioned, what to watch.
Problem & market · 16:9

The "unlimited era" is over

Flat-rate AI subscriptions can no longer absorb how people actually use frontier models. The shift to metered, rolling-window access is now the norm rather than the exception — and it's structural, not stingy.

ProviderLimit structure (2026)What users feel
ChatGPTRolling windows: Plus/Go ≈ 160 GPT-5.5 messages / 3h; Free ≈ 10 / 5h then silent downgrade to a "mini" model. Reasoning models carry separate weekly & monthly caps.Silent downgrades; unclear which cap you're hitting.
ClaudeTwo limits at once — a 5-hour rolling window and a separate weekly cap. May 2026 doubled the 5h limits & removed peak throttling, but left the weekly cap untouched. Opus-class can drain quota 3–5× faster.Heavy users "run out by Wednesday."
Root causeDemand is outrunning GPU supply. After a surge of new users in early 2026, inference capacity couldn't scale as fast as sign-ups.Tighter limits, slower responses, downgraded models.

The real pain is unpredictability, not scarcity. The limit itself is tolerable. What breaks the experience is that it's opaque and arbitrary from the user's seat — the same task costs wildly different amounts depending on length, tool use, artifacts, and model. Users can't see the meter, can't budget against it, and get cut off mid-thought.

Pain 01Acute
Lost momentum.
The wall hits at the worst moment — mid-task, with no graceful handoff. "It feels like the tool quit on you."
Pain 02Acute
No forward planning.
Users can't stage work against a reset they can't clearly see. The meter is invisible, so nothing can be budgeted against it.
Pain 03Acute
Dead time.
Between hitting the wall and the reset, productive intent evaporates. The time until reset is pure loss.
Tailwind · local LLMsEnabler
A free "good enough" bridge is now viable.
Open-weight local models crossed the "good enough" threshold in 2026 — laptop-friendly Qwen sizes (0.6B–8B) run in a browser/desktop runtime with no key and no bill, making a free first-draft engine economically and technically real.
3×
Year-over-year growth in local-model adoption among developers.
Local-LLM adoption reporting, 2026
~85%
Of flagship-level quality from mid-range local models on everyday writing, drafting, summarization.
2026 estimates
0.6–8B
Qwen3 sizes that run in a laptop browser/desktop runtime — no key, offline.
Qwen3 model card
$0
API cost for the local bridge — full data privacy, fully offline-capable.
Local inference
The competitive gap · the wedge
Competitors schedule notifications. AutoQ schedules execution — timed precisely to the reset gate — and keeps the user moving with a free local model while the clock runs down.
OpenAI's June 17, 2026 Scheduled Tasks relaunch validates clock-driven AI, but it is a reminder/monitoring tool (≈ once per hour, may auto-pause). It does not stage blocked work against the user's own reset window, and it does not bridge to a free model in the meantime.
Persona walkthrough · 16:9

Who hits the wall

North-American builders and knowledge workers who pay for frontier AI and hit its limits several times a week. They want to get unblocked — not to optimize prompts as a hobby.

Solo builder who hits the 5-hour or weekly wall 2–4× a week, usually mid-afternoon.
Primary · Maya · 29 · Austin, TX

Indie developer / early-stage founder. Works in long, bursty sessions and values getting unblocked over saving pennies.

Plan
One premium AI plan (Claude Pro or ChatGPT Plus, ~$20/mo), occasionally a second.
Usage
Long, bursty sessions; hits the wall 2–4×/week, usually mid-afternoon.
Attitude
Not a prompt-engineering hobbyist — wants output, not optimization homework.
Comfort
Happy to install a lightweight desktop/browser tool; unblocking > saving pennies.
Async worker who batches research & drafting and lines up work overnight.
Secondary · Deic · 35 · Toronto

PM who batches research and drafting tasks and works across time zones. Wants to "reserve while I sleep" and wake to finished drafts.

Plan
Premium AI plan; heavy batch usage across the workday and off-hours.
Usage
Batches tasks; works across time zones; overnight staging is ideal.
Attitude
Planning-minded; comfortable staging work against a schedule.
Job
"Reserve while I sleep" — wake up to completed drafts.
Jobs to be done
When I hit a limit mid-task, help me not lose momentum or the thread of what I was doing. When I know I'll be away, let me stage work so it runs the instant I have capacity again. While I'm blocked, give me a "good enough" draft now so I'm reviewing instead of waiting.
Flow & data-model demo · 16:9

The reservation, step by step

The end-to-end happy path. Each ★ strength marks a moment where AutoQ creates value today's experience does not — the As-Is → To-Be resolution points, made concrete.

Entry · the wall
Usage limit hit — momentum stops
Maya is deep in a Claude session refining a product-spec draft and sends a long prompt. The native tool returns "You've reached your limit — try again after 4:30 PM." Dead end.
★1 · Graceful interception
Instead of a dead-end error, AutoQ detects the limit event and surfaces a non-blocking prompt: "Hit the wall? Reserve this and keep moving." The failure state becomes an entry point.
Resolves: abrupt, momentum-killing cutoff.
Capture · reserve the work
The blocked prompt becomes a task queue
Maya clicks Reserve. AutoQ captures the blocked prompt plus lightweight context (the last few turns, the target model, file references) and decomposes it into discrete sub-tasks it can run in sequence — outline → draft A → draft B → consolidate.
★2 · Decomposition into a schedulable unit
The vague "big ask" becomes an ordered, resumable queue — the thing that makes reservation and partial local drafting possible.
Resolves: all-or-nothing prompting; no way to stage work.
Schedule · target the reset gate
Timed to the user's actual window
AutoQ reads the reset time and shows a countdown with a proposed run: "Auto-runs at 4:30 PM on Claude." Maya confirms or edits ("run overnight," "run on my next weekly reset") and can queue multiple reservations.
★3 · Reset-gate targeting
Work is scheduled to the user's actual window, not an arbitrary hourly cron. The moment capacity returns, the queue fires — nothing wasted.
Resolves: dead time between wall and reset; planning against an opaque meter.
Bridge · keep moving now
A free local draft while the clock runs
While the countdown runs, AutoQ offers a first pass on a free local Qwen web app — a rough draft in seconds, no key, no cost, offline-capable. Maya reviews and tightens it, and her edits feed back into the reserved prompt so the scheduled premium run starts from a sharper brief.
★4 · Free bridge that compounds
Waiting time becomes drafting time. The local pass improves the input for the premium run — the eventual paid tokens are spent on a better-scoped task.
Resolves: zero productivity while blocked; premium tokens wasted on under-specified prompts.
Fire · the reset executes
Set-and-forget completion
At 4:30 PM the window resets. AutoQ auto-executes the queued tasks on the premium model, in order, and notifies: "Your spec draft is ready." Maya opens a finished, premium-quality result that already incorporates her local-pass refinements.
★5 · Set-and-forget completion
She reserved work and walked away; AutoQ delivered it the instant capacity allowed. The wall became a reservation with a guaranteed seat.
Resolves: user must babysit the reset and re-submit manually.
Loop · learn and pre-empt
From rescue to planning
AutoQ logs the run (tokens estimated, model, reset window) and, over time, predicts when Maya tends to hit walls — proactively suggesting reservations before she's blocked.
★6 · Predictive pre-emption
The service shifts from reactive rescue to proactive planning — the stickiness layer that makes AutoQ a habit, not a fire extinguisher.
Resolves: recurring, reactive frustration.
Flow at a glancewall → reservation → reset → completion
Working in premium AI Usage limit hit Capture prompt + ctx Decompose → queue Read reset → countdown Reset → auto-executepremium model, in order Bridge · free local Qwendraft now · no key · offline Completion notification Log + predict next wall ★1 Graceful interception ★2 Schedulable unit ★3 Reset-gate targeting ★4 Compounding bridge ★5 Set-and-forget ★6 Predictive pre-emption edits feed back → sharper brief

The data model

Reservation → Task → (BridgeRun | Execution) is the spine of the product. Decomposition (★2) is what lets the same task be drafted for free locally and executed premium at reset — the two-track mechanic no competitor's "scheduled task" replicates. Attributes are abbreviated below; the full relationships are tabled underneath.

Userentity user_id PK email · plan_type · timezone settings_id FK created_at UsageWindowentity window_id PK user_id · provider_id FK window_type · reset_at est_remaining_quota last_hit_at UsagePredictionfeeds ★6 prediction_id PK user_id · provider_id FK predicted_wall_at confidence · suggested_action AIProviderentity provider_id PK name Claude · ChatGPT · … model_options[] limit_type rolling_5h · weekly default_reset_logic Reservation★ core object reservation_id PK user_id · provider_id FK original_prompt · context_snapshot status draft·scheduled·bridging·running·done scheduled_for · created_at Notificationentity notification_id PK user_id · reservation_id FK type reserved·reset_soon·completed·failed sent_at · read Tasksub-unit task_id PK reservation_id FK sequence_order · instruction depends_on task_id | null status · output_ref BridgeRunfree local pass bridge_run_id PK task_id FK local_model qwen3:1.7b draft_output · user_edits fed_back · ran_at Executionpremium run execution_id PK task_id · provider_id FK model_used · tokens_estimated result_output executed_at · success 1·N 1·N 1·N makes 1·N 1·N decomposes 1·N 1·0..N 1·0..1 provider · 1·N execution
FromCardinalityToMeaning
User1 — NReservationA user makes many reservations.
User1 — NUsageWindowOne window per (user, provider) limit type.
AIProvider1 — NReservationEach reservation targets one provider.
Reservation1 — NTaskA reservation decomposes into an ordered set of tasks.
Task1 — 0..NBridgeRunA task may get zero or more free local draft passes.
Task1 — 0..1ExecutionA task resolves to one premium execution at reset.
Reservation1 — NNotificationStatus changes emit notifications.
User1 — NUsagePredictionPredictions drive proactive suggestions.
UsageWindow1 — NUsagePredictionPredictions derive from observed window behavior.
Design note
Reservation → Task → (BridgeRun | Execution) is the spine. Decomposition is what lets the same task be drafted for free locally and executed premium at reset — the two-track mechanic that no competitor's "scheduled task" replicates.

Who's in the system

Concentric rings by proximity to the core mechanic: primary actors trade value directly with AutoQ; secondary are the platforms and infrastructure it rides on; tertiary set the context and the rules. Four categories divide the field.

1 · Primary 2 · Secondary 3 · Tertiary Users AI Platforms Local & Infra Governance & Distribution Solo builder Async worker Small teams Enterprise admins Prompt communities Paid AI account Anthropic · Claude OpenAI · ChatGPT Google · Gemini Emerging providers GPU / compute supply Qwen local model Ollama · llama.cpp WebGPU browser OpenRouter Hugging Face Device makers Platform ToS / policy App stores Privacy regulators Competitor · Scheduled Tasks AutoQ scheduling layer
1 · Primary — direct value exchange2 · Secondary — ecosystem & dependencies3 · Tertiary — context & governance

Categories. Users (the demand), AI Platforms (the paid windows AutoQ schedules around), Local & Infra (the free bridge and its runtime), and Governance & Distribution (the rules and channels). The nearer the ring, the more directly a stakeholder shapes the reservation loop — users, the paid account, and the local Qwen model sit at the core; regulators and competitors sit at the edge.

Metrics & strategy · 16:9

How we know it works

One number captures the promise; a funnel of supporting metrics explains why it moves.

North Star Metric
Reserved Tasks Successfully Completed at Reset — per active user, per week.
It only moves when the whole loop — interception → reservation → reset execution — actually delivers work the user would otherwise have lost.
Activation
Limit-hit → Reservation rate
Of walls hit, % that become reservations — is ★1 interception working?
Time-to-first-reservation
After install.
Engagement
Bridge-run adoption
% of reservations that use the free local draft — measures ★4 pull.
Feedback-loop rate
% of bridge drafts whose edits feed back into the premium run — the compounding-value signal.
Avg. tasks per reservation
Decomposition depth.
Reliability (trust)
Reset-execution success rate
% of scheduled runs that fire correctly at reset — the trust backbone; must be near-100%.
Median reset-drift
Seconds between actual reset and execution.
Retention
W1 / W4 retention
Of users who completed ≥ 1 reserved task.
Proactive-suggestion acceptance
% of predicted-wall suggestions acted on — measures ★6 stickiness.
Business / value
Recovered productive time
Dead-time-blocked hours converted to draft/complete, per user per week.
Premium-token efficiency lift
Token cost per task, with vs. without a fed-back bridge draft.
Willingness-to-pay proxy
Reservations per paying vs. free-tier user.

From dead end to reservation

Every friction of today's experience maps to a resolution on the flow.

DimensionAs-Is (today)To-Be (with AutoQ)Resolved
Hitting the limitHard error, mid-task; momentum dies.Non-blocking interception → "reserve & keep moving."★1
The blocked promptLost or manually re-pasted later.Captured, decomposed into a resumable task queue.★2
The reset windowOpaque, arbitrary; user must remember & re-submit.Read, counted down, auto-targeted for execution.★3 · ★5
Wait timeDead time — pure loss of productivity.Draft time — free local Qwen pass produces a result now.★4
Prompt qualityPremium tokens spent on under-specified asks.Local pass sharpens the brief; premium run starts better-scoped.★4
Getting the resultUser babysits the reset, re-runs manually.Set-and-forget: fires at reset, notifies on completion.★5
Recurring wallsReactive frustration every time.Predictive suggestions pre-empt the next wall.★6
The core transformation
As-Is, a usage limit is a dead end. To-Be, the same limit becomes a reservation — work staged, kept warm on a free model, and delivered the instant capacity returns.
AutoQ converts the industry's biggest UX friction — opaque, momentum-killing limits — into a planning surface the user controls.

Why this wins

AutoQ monetizes a friction that grows as the underlying trend deepens — and sits in an adjacent space competitors have validated but not occupied.

Moat 01Trend
Rides an irreversible trend.
Metered, windowed access is a supply/demand reality as inference demand outpaces GPU capacity. AutoQ monetizes the friction that trend creates, so its relevance grows as limits tighten.
Moat 02Neutral
Platform-agnostic layer.
AutoQ sits above Claude, ChatGPT, Gemini — it doesn't compete with the models, it makes each less frustrating. Users consolidate multi-tool usage-management into one layer.
Moat 03Differentiator
Two-track cost structure.
The free local bridge (Qwen) is near-zero marginal cost, yet it keeps users engaged during the wait and improves premium output quality. No competitor's "scheduled task" pairs reset-timed execution with a free drafting bridge.
Moat 04Timing
Category timing.
OpenAI's June 2026 Scheduled Tasks relaunch proves the market is moving to clock-driven AI — but it stopped at reminders. AutoQ occupies the unclaimed adjacent space: clock-driven execution against the user's own limits.
Business & positioning

Acute, recurring pain justifies a subscription.

Users hit walls multiple times a week; a tool that reliably recovers lost work justifies $5–15/mo — well below the AI plans it protects. Retention is structural: once reservations and predictions accumulate, AutoQ becomes the user's AI "control tower," and switching cost rises with every logged window (★6). Expansion runs solo builders → teams (shared queues, follow-the-sun scheduling) → enterprise (org-level usage governance and cost visibility).

Risks to watch
Platform dependency
If a provider ships native reservations, the edge narrows — mitigated by staying cross-platform; value grows with the number of tools managed.
ToS / automation posture
Auto-executing at reset must respect each provider's terms — positioned as user-initiated, user-owned scheduling, not circumvention.
Bridge-model quality ceiling
Local drafts are ~85% quality, not frontier — framed as "draft now, perfect at reset," never as a replacement.

Competitors schedule notifications. AutoQ schedules execution — timed to the reset gate, bridged by a free local model — turning the industry's biggest UX friction into a planning surface the user controls.

AutoQ — UX Strategy Document · Product Design, AI Token Tool

Prepared as a UX strategy artifact. Market figures reflect publicly reported 2026 usage-limit and local-LLM data (ChatGPT/Claude usage documentation and tracking, local-LLM adoption reporting, and OpenAI's June 17, 2026 Scheduled Tasks relaunch) and should be re-verified before external publication, as AI pricing and limits change frequently. Personas are composite profiles, not real individuals.