Persistence¶
CascadeUI state is ephemeral by default. Nothing reaches disk until you opt a slot in. Two independent namespaces cover the two things that can survive a restart, each served by a backend that declares its capabilities up front:
| Namespace | Contents | Config class |
|---|---|---|
registry |
PersistentView reattach rows (one row per persistence_key) |
RegistryPersistence |
application |
Reducer slots opted in via persistent_slots or access_slot(..., persistent=True) |
ApplicationPersistence |
PersistenceMiddleware owns the full startup pipeline: wire backends to
namespaces, apply migrations, block until rehydrate completes, install the
write-through dispatch hook, and reattach persistent views when a bot is
supplied. Install it through setup_middleware, which awaits each
middleware's initialize(store) method in order.
Setup¶
Construct PersistenceMiddleware once in your bot's setup_hook,
after loading your cogs:
from discord.ext import commands
from cascadeui import PersistenceMiddleware, setup_middleware
from cascadeui.persistence import SQLiteBackend
class MyBot(commands.Bot):
async def setup_hook(self):
# Load cogs first so PersistentView subclasses register themselves
# via __init_subclass__ before the middleware looks them up
await self.load_extension("cogs.dashboard")
await self.load_extension("cogs.counter")
# One backend covers both namespaces (shorthand form)
await setup_middleware(
PersistenceMiddleware(backend=SQLiteBackend("cascadeui.db"), bot=self),
)
Cog loading order matters
The persistence middleware must initialize after every cog that
defines a PersistentView subclass is imported. Python's import
machinery populates the class registry via __init_subclass__;
initializing against an empty registry lands every surviving panel in
the skipped bucket, dead until a reattach() re-drive or the next
restart picks it up.
Cogs do not install middleware
Middleware install belongs to the bot author, not a cog. A cog that
called setup_middleware(...) inside its own setup(bot) would
silently mutate the bot's store without the author's consent.
Declare the dependency in the cog's docstring instead and let the
bot's setup_hook satisfy it.
Zero-config construction is allowed
PersistenceMiddleware() with no arguments defaults to
SQLiteBackend("cascadeui.db") for both namespaces. The optional
aiosqlite dependency is required; without it initialize raises
PersistenceInitError with an install hint. Zero-config is
ephemeral-in-practice because no slots are opted in yet -- install
it, then opt slots in with persistent_slots = ("name",) on your
view class (or access_slot(..., persistent=True) from a reducer)
as you need them.
Per-namespace configuration¶
The shorthand backend= fills any namespace that was not given an explicit
config. Passing a namespace config overrides the shorthand for that
namespace:
from cascadeui import PersistenceMiddleware, setup_middleware
from cascadeui.persistence import (
InMemoryBackend,
SQLiteBackend,
RegistryPersistence,
ApplicationPersistence,
SlotPolicy,
)
await setup_middleware(
PersistenceMiddleware(
# Shorthand: any namespace without an explicit config uses this
backend=SQLiteBackend("cascadeui.db"),
# Application slots use a separate SQLite file with per-slot policies.
# Slots not listed here still default to ephemeral -- list only the
# slots you want durable.
application=ApplicationPersistence(
backend=SQLiteBackend("application.db"),
slots={
"user_preferences": SlotPolicy(persistent=True),
"search_cache": SlotPolicy(persistent=True, ttl_days=7),
},
),
bot=bot,
),
)
Pass backend=None inside a specific namespace config to opt that namespace
out entirely. The shorthand does not override an explicit config, so this
is the canonical way to turn one namespace off while leaving the other on:
await setup_middleware(
PersistenceMiddleware(
backend=SQLiteBackend("cascadeui.db"),
application=ApplicationPersistence(backend=None), # registry only
bot=bot,
),
)
Data-only vs reattach¶
bot= is optional. Without it, the middleware initializes backends and
rehydrates state, but skips persistent-view reattach:
# Data-only: state survives restart, but PersistentView panels do not
# re-attach to their original messages
await setup_middleware(
PersistenceMiddleware(backend=SQLiteBackend("cascadeui.db")),
)
Data-only mode still restores any slots marked persistent=True.
Full mode logs a reattach summary with five buckets, and callers that
need the structured result can await reattach_persistent_views()
directly on the manager:
summary = await store.persistence_manager.reattach_persistent_views()
# {"restored": [...], "skipped": [...], "failed": [...], "removed": [...], "unreachable": [...]}
restored: view reattached successfully.skipped: view class not imported, or kwargs migrator missing. Row stays on disk so the next restart retries.failed: construction or a kwargs migrator raised during reattach, or a 404-verdicted row could not be confirmed or pruned because the backend raised. (on_restoreruns later, after the bot is ready; its failures are logged there, not reflected in this bucket.) Row stays on disk; the next pass retries it.removed: channel or message returned a definitive 404 (discord.NotFound), re-confirmed against the live row. Row removed, andREGISTRY_PRUNEDdispatched withreason="gone"andsource="reattach".unreachable: channel or message could not be fetched for a transient reason (Forbidden,RateLimited,HTTPException, a transport failure, or a non-messageable channel), or the row was re-posted under its key mid-pass and now points at a message the pass never fetched. The row is left on disk so a clean restart retries; nothing is pruned. Do not reconcile external records from this bucket: the panel may still exist.
Rows that stay unreachable¶
Keeping an unreachable row is what stops a permission change during startup
from deleting a live panel, and it has a cost: a channel the bot will never
see again is re-fetched at every boot, forever. Nothing in a Forbidden
distinguishes "not right now" from "not ever".
So the first failed fetch stamps the row with first_unreachable_at, and the
stamp clears the moment a fetch succeeds. It records the first failure, not
the latest, so the age keeps growing across boots rather than resetting.
mgr = store.persistence_manager
mgr.unreachable_since
# {"tickets:panel": 1754300000}
await mgr.prune_unreachable(older_than_days=30)
# {"pruned": [...], "gone": [...], "recovered": [...], "kept": [...]}
prune_unreachable re-checks every candidate against Discord before deleting
anything. A row that answers is kept and its stamp cleared, however old that
stamp was; a row returning a definitive 404 goes regardless of age; the rest
are deleted only if the stamp predates the cutoff. A stamp on its own is one
observation, and a host that simply has not restarted for a month carries a
month-old stamp from a single failure -- the second look is what turns it into
evidence. The deletions dispatch REGISTRY_PRUNED with
source="prune_unreachable": one with reason="gone" for the rows whose channel
or message no longer exists (listed in gone as well), and one with
reason="unreachable" for the rows that aged out.
A row whose key a live panel in this process holds is kept without a fetch: the
panel owns the key, and a failed fetch says nothing about whether its row should
go. It raises ValueError for a negative or non-integer cutoff and RuntimeError
when the middleware was built without bot=, since there is nothing to
re-verify against. /cascadeui unreachable lists the backlog with ages, and
takes an optional prune_older_than_days to run the same prune from Discord.
To run it without a caller, pass prune_unreachable_after_days=:
await setup_middleware(
PersistenceMiddleware(
backend=SQLiteBackend("cascadeui.db"),
bot=self,
prune_unreachable_after_days=30,
),
)
The first prune runs once the gateway is ready, never during setup_hook, so
every reattach() you issue there has finished, and a fetch that fails with the
gateway up points at the channel rather than the host. It then repeats daily and
stops when the manager closes. It never overlaps a reattach pass. The setting is
off by default, takes a positive integer, and needs bot=; a zero cutoff is
refused because an automatic run would then delete a row on a single failed
fetch. Several processes sharing one registry each run their own prune, with no
coordination between them.
The clear is keyed on the fetch succeeding, not on the view reconstructing.
A row whose class raises during construction lands in failed, but its
message was reachable, so its stamp clears and it is never a prune candidate.
The middleware runs reattach_persistent_views() inside setup_middleware,
after cogs are loaded but before on_ready fires. A REGISTRY_PRUNED
subscription registered in a cog's setup(bot) is already in place and does
observe the action. Code that subscribes only in on_ready or later misses it
(the action dispatches once at startup, with no replay). To reconcile records
kept outside the registry, read the stashed summary after startup instead:
from cascadeui import get_store
@bot.event
async def on_ready():
summary = get_store().persistence_manager.total_reattach_summary
for key in summary["removed"]:
... # clear your own row for this persistence_key
total_reattach_summary covers every pass, with each key under the outcome the
most recent pass gave it. Read it rather than last_reattach_summary, which
holds one pass: a re-drive
replaces it before on_ready runs, and it cannot re-report a removal, because
the row the first pass deleted is no longer there to verdict. Reconcile only
from removed (a definitive 404); a key in unreachable may still exist and
should be left alone. The keys list on the REGISTRY_PRUNED action with
reason="gone" carries the same removed data for a consumer that
subscribes before setup_middleware. A removal made after a subscriber
registered, at boot or by a re-drive, reaches it both ways, so act on
reattach removals from one of the two: the action's source is
"reattach" for those, and "prune_unreachable" or "prune_registry" for the
deletions that come later.
Re-driving reattach after a runtime cog load¶
With the canonical setup_hook order above (cogs first, then setup_middleware),
the reattach pass inside setup_middleware imports every PersistentView
subclass before it runs, so all of them attach. That order needs no reattach()
call.
reattach() is the escape hatch for a persistent-view cog loaded after
setup_middleware has already run: a hot-reloaded extension, or a cog loaded
lazily at runtime. Its class was not imported during the initial pass, so its
posted messages land in the skipped bucket and stay dead until the next
restart. Call reattach() once after the runtime load to attach them:
async def reload_feature(self):
# A persistent-view cog loaded well after setup_hook finished.
await self.load_extension("cogs.new_dashboard")
await get_store().persistence_manager.reattach()
reattach() is idempotent: panels already attached on a prior pass are
skipped (no re-fetch, no double registration), and transiently unreachable /
failed rows are retried. A key a live panel in this process already holds is
skipped too, even when its row never attached at startup, since re-driving it
would attach a second instance to the message that panel owns. It returns the
same five-bucket summary as
reattach_persistent_views(), covering only the rows it processed. Each pass
also re-registers every DynamicPersistentButton subclass with the bot, so a
dynamic button defined in the late-loaded cog routes clicks too.
Shutdown¶
Registry writes reach the backend at once. Application-slot writes are batched: a change is written about two seconds after the last one, and no later than ten seconds after it was made. Closing persistence writes whatever is still batched and then closes the backends.
With bot=, persistence closes when the bot does, whether bot.run() stops
or your code calls await bot.close(), from a shutdown command or anywhere
else. The bot's own close() runs first, including an override on your bot
class, so a cog that saves state while it unloads is written too.
On Linux and macOS, systemd, docker stop, and most hosting panels stop a
program with SIGTERM, which discord.py leaves at its default: the process
dies on the spot. So with bot=, the first SIGTERM closes the bot instead,
which ends a program that runs the bot as its main task (bot.run(), or
bot.start() in async with bot). A program running other work beside the
bot keeps running, and a second SIGTERM ends it at once. Once the bot has
closed, with nothing left for SIGTERM to close, a SIGTERM ends the process
as it would without CascadeUI. A bot started again in the same process after
a SIGTERM (a retry loop that builds a new bot) ends the process
instead of running on until the process manager kills it, since the process
was asked to stop; a change made since the close is written first. A
SIGTERM handler of your own, installed with signal.signal() or
loop.add_signal_handler() before or after setup_middleware(), takes
precedence. On Windows, a process
ended from outside (Task Manager, most process managers) runs no cleanup at
all, so stop a bot there with Ctrl+C or await bot.close().
Without bot=, the library has no way to know the process is stopping, so
close it yourself before the event loop ends:
A process that skips this loses the batched writes, and on SQLiteBackend it
does not exit at all: the database runs on a worker thread that keeps the
interpreter alive until the backend closes. Closing twice is harmless, even
from two tasks at once. Each backend gets up to ten seconds to close; one
that takes longer is logged and persistence closes without it, so the bot's
close() still returns.
Restarting in the same process¶
A dropped gateway connection does not close the bot: with the default
reconnect=True, discord.py resumes it or connects again inside
connect(), and every view stays live. It closes the bot itself only on a
close code it cannot recover from, such as a rejected token or an invalid
shard.
A closed bot cannot log in or connect again, even after clear(), so a
restart in the same process builds a new bot object, and its setup_hook
runs setup_middleware() again. That passes the new bot to the installed
middleware and reopens persistence, which closes with the new bot from then
on. A change made while persistence is closed is held in memory and written
when it reopens, and the first one logs a warning, since it is lost if the
process exits instead.
The views a bot leaves are handled as a process restart would handle them,
with or without persistence. When the bot closes, each view sent through it
stops: its timeout never fires, its instance slot is freed, code awaiting
its wait() is cancelled as it would be when the process ends (a later
wait() on it raises asyncio.CancelledError too), and nothing is edited,
since the client it was sent through has closed. They leave the state with
their sessions at the close, and no action is dispatched for them: no
VIEW_DESTROYED reaches a hook or a middleware, as none would after a
process restart. The release runs inside the bot's close(), which is
wrapped when setup_middleware() runs or the bot sends its first view; a
reference to bot.close taken before that, or a call to the class's own
close(), skips the release, and persistence's close with it.
A persistent panel among them keeps its registration and is restored
through the new bot when its setup_hook runs setup_middleware(), before
the bot connects, so the panel is back in place by on_ready. A bot that
closes while its setup_hook is still restoring leaves the panels it had
not restored registered, and the next bot restores them. Code that kept
a reference to the old panel reaches the restored one through
current_view; on any other view from
before the restart, refresh() and exit() do nothing.
Backends¶
Built-in backends¶
The library ships three backends:
| Backend | Import | Capabilities |
|---|---|---|
InMemoryBackend |
cascadeui.persistence |
KV, RELATIONAL, TTL_INDEX, SCHEMA_META, OPEN_ROWS |
SQLiteBackend |
cascadeui.persistence (requires aiosqlite) |
KV, RELATIONAL, TTL_INDEX, SCHEMA_META, RAW_SQL |
PostgresBackend |
cascadeui.persistence (requires asyncpg) |
KV, RELATIONAL, TTL_INDEX, SCHEMA_META, RAW_SQL |
InMemoryBackend is the reference implementation. It matches the Protocol
exactly and is useful for tests. SQLiteBackend is the recommended
single-process production default: WAL mode for concurrent reads,
ON CONFLICT upsert, NULL-safe TTL prune, and LIKE-ESCAPE safe scan.
PostgresBackend adds cross-process coordination via LISTEN/NOTIFY
and is the right choice for multi-process deployments.
The tables a SQL backend creates¶
Both SQL backends create four tables and two indexes:
| Table | Holds |
|---|---|
cascadeui_persistent_views |
One reattach row per persistence_key |
cascadeui_application_slots |
One row per persistent application slot |
cascadeui_schema |
Per-table schema version, read by the migrator |
cascadeui_kv |
The namespaced key-value surface |
All four carry the cascadeui_ prefix, so a scope written against that
pattern (a filtered backup, a grant, an audit query) reaches every table the
library owns, and the names are distinctive enough not to collide with a
consumer's own schema when both share a database.
A database created before the registry and slots tables carried the prefix
is renamed in place: apply_migrations finds persistent_views or
application_slots, confirms the table carries the library's columns, and
renames it (indexes and schema-version record included) in one
transaction before schema versions resolve. A same-named table without
the library's columns is a consumer's own and is never touched; when both
the old and new names exist and both hold rows, nothing is renamed, the
library uses the new name, and a WARNING names both tables so an operator
can reconcile.
Pass table_prefix= to move every table and index the backend owns, which
keeps two CascadeUI deployments sharing one database apart:
The prefix applies to all four tables and both indexes. It is empty by default. Changing it on a database that already holds rows points the backend at a fresh, empty set of tables; the old rows stay where they are.
A prefix that is not lower-case letters, digits, and underscores (starting
with a letter or underscore) raises ValueError at construction. The
prefix reaches SQL through two paths that quote identifiers differently:
the DDL and the migrator interpolate it as raw text, while a backend's own
row and key-value paths quote what they build. PostgreSQL folds an unquoted
identifier to lower case and preserves a quoted one, so a prefix carrying
a capital would create one table and address another, splitting a
deployment's data across two that both look right on their own. SQLite
matches identifiers case-insensitively and is not affected, but the same
prefix moved to PostgreSQL later is, so both backends refuse it.
Inspecting the database directly
CascadeUI partitions its state across dedicated tables, not one blob.
cascadeui_persistent_views holds the PersistentView registry -- one
row per posted panel, written through immediately on registration.
cascadeui_application_slots holds persisted application and scoped
slots. cascadeui_kv is a generic key-value surface for the KV Protocol
methods and does not hold view registry rows. To confirm a panel's
row exists, query cascadeui_persistent_views, not cascadeui_kv.
from cascadeui import PersistenceMiddleware, setup_middleware
from cascadeui.persistence import SQLiteBackend
await setup_middleware(
PersistenceMiddleware(backend=SQLiteBackend("cascadeui.db"), bot=bot),
)
PostgreSQL backend¶
PostgresBackend ships full Protocol surface against PostgreSQL via
asyncpg. Install the optional dependency:
Configure with a connection string:
import os
from cascadeui import PersistenceMiddleware, setup_middleware
from cascadeui.persistence import PostgresBackend
backend = PostgresBackend(dsn=os.environ["CASCADEUI_DATABASE_URL"])
await setup_middleware(
PersistenceMiddleware(backend=backend, bot=bot),
)
The dsn accepts the standard libpq URL format. Production deployments
use sslmode=verify-full for full TLS certificate verification:
Required database privileges¶
The CascadeUI database user needs minimal GRANTs:
GRANT CONNECT ON DATABASE cascadeui TO cascadeui_app;
GRANT USAGE ON SCHEMA public TO cascadeui_app;
GRANT SELECT, INSERT, UPDATE, DELETE
ON cascadeui_persistent_views, cascadeui_application_slots,
cascadeui_kv, cascadeui_schema
TO cascadeui_app;
Schema migrations need more than this
A version bump that alters a table's columns (see Migrations)
runs ALTER TABLE during apply_migrations(), and PostgreSQL grants no
privilege for that short of table ownership -- the DML grants above are not
enough. Under a locked-down cascadeui_app role the migration fails and the
bot does not boot, raising PersistenceSchemaError naming the table, the
version step, and the driver's own error. The on-disk version is left
unchanged, so the migration simply runs on the next start once the
privilege is there. Run it once as the table owner (or a role with ALTER
rights), then resume running the bot as cascadeui_app.
SQLite has no equivalent step: file-level write access covers ALTER TABLE
the same as any other write.
A database still carrying the pre-rename table names meets this wall one
step earlier. The in-place rename described above runs before any version
step, and ALTER TABLE ... RENAME TO needs the same ownership, so on a
locked-down role it is the first statement to fail. It raises
PersistenceSchemaError naming both table names and the same remedy, and
the rename commits as one transaction, so the database is unchanged and it
retries once the privilege is there. A role that cannot see the old table
at all skips the rename silently and reaches the version step as before.
Cross-process invalidation¶
PostgresBackend uses LISTEN/NOTIFY to broadcast slot invalidations
to other CascadeUI processes connected to the same database. Each worker
hears about another worker's writes through the callback below. The listener
connection sits outside the connection pool (LISTEN registrations are
session-scoped per the PostgreSQL contract) and auto-reconnects on drop.
Each worker keeps its own copy of a slot
The library does not reload a slot into memory on its own, so a worker that ignores the notification keeps serving its own copy. A worker also skips a write that would store what it last stored: once another worker has changed a slot, setting the slot on this worker back to the value it last stored writes nothing, and the other worker's value stays on disk. The expiry sweep reads this worker's own record too, so a slot whose expiry, as this worker last recorded it, has passed leaves its memory even when another worker has written the row since. The row itself stays.
The channel carries table_prefix when one is set, so two deployments
sharing a database stay off each other's bus the same way they stay out of
each other's tables. Workers of the same deployment share a prefix and so
still hear about each other's writes, which is the point. A notification names the
logical namespace and key, and the key is empty when the payload would
exceed PostgreSQL's 8000-byte limit, meaning "drop this whole namespace" --
so a shared channel would turn one neighbour's large write into a full cache
flush here.
Register a per-process callback to consume the invalidation stream:
def on_invalidate(namespace: str, key: str) -> None:
# Drop any local cache entry keyed by (namespace, key)
cache.invalidate(namespace, key)
backend.set_invalidation_callback(on_invalidate)
Pool tuning¶
Library defaults: min_size=2, max_size=10, statement_cache_size=1024.
Override via pool_kwargs:
pgbouncer compatibility¶
asyncpg's prepared-statement cache requires session-mode pooling.
Operators running pgbouncer in transaction or statement mode set
statement_cache_size=0:
Custom tables and raw SQL¶
CascadeUI's namespace API (row_upsert / row_select / kv_*)
covers the common cases. Backends that declare Capability.RAW_SQL
expose a raw-SQL escape hatch for everything else: domain tables in the
same database, vendor-specific features (PostgreSQL JSONB GIN queries,
SQLite FTS5), custom indexes, and ad-hoc analytics queries.
Three patterns to choose from:
Pattern A: Separate database (recommended for unrelated domain data)¶
If your bot's domain data has nothing to do with CascadeUI state, keep them apart. Open your own database connection, run your own migrations, manage your own schema. CascadeUI's persistence layer stays focused; your domain code stays portable across CascadeUI versions.
Pattern B: KV escape hatch (recommended for opaque blobs)¶
Use the existing KV surface with a custom namespace:
backend = store.persistence_manager.application.backend
if backend is None:
raise RuntimeError("Application namespace has no backend configured")
await backend.kv_write("ticket_threads", "guild:42:thread:99", json.dumps(data).encode())
ticket_data = await backend.kv_read("ticket_threads", "guild:42:thread:99")
async for key, value in backend.kv_scan("ticket_threads", prefix="guild:42:"):
...
Zero schema work, full integration with the persistence pipeline, cross-backend portable. Limited to opaque bytes payloads (no relational queries, no joins).
When backend may be None
store.persistence_manager.application.backend is None when the
application namespace was opted out (application=ApplicationPersistence(backend=None)).
Patterns B and C apply only when the application namespace is
backed; guard with the is None check above before reaching for
the escape hatch.
Pattern C: Raw SQL escape (for SQL-rich data co-located with CascadeUI)¶
Backends declaring Capability.RAW_SQL expose four query methods plus
an explicit transaction primitive. Check the capability first:
from cascadeui.persistence import Capability
backend = store.persistence_manager.application.backend
if backend is None:
raise RuntimeError("Application namespace has no backend configured")
if Capability.RAW_SQL not in backend.capabilities:
raise RuntimeError("Backend does not support raw SQL")
Each backend reports its parameter syntax through placeholder_style
(PEP 249 paramstyle): "qmark" for SQLite, "numeric" for PostgreSQL.
Portable code adapts at write time:
ph = backend.placeholder_style
# Create a custom table
await backend.execute("""
CREATE TABLE IF NOT EXISTS tickets (
id INTEGER PRIMARY KEY,
user_id BIGINT NOT NULL,
content TEXT NOT NULL,
created_at BIGINT NOT NULL
)
""")
# Insert with portable placeholder formatting
sql = (
"INSERT INTO tickets VALUES (?, ?, ?, ?)" if ph == "qmark"
else "INSERT INTO tickets VALUES ($1, $2, $3, $4)"
)
await backend.execute(sql, 1, user_id, content, int(time.time()))
# Query
rows = await backend.fetch(
"SELECT * FROM tickets WHERE user_id = ?" if ph == "qmark"
else "SELECT * FROM tickets WHERE user_id = $1",
user_id,
)
# Single-row lookup
row = await backend.fetch_one(
"SELECT * FROM tickets WHERE id = ?" if ph == "qmark"
else "SELECT * FROM tickets WHERE id = $1",
ticket_id,
)
if row is None:
raise LookupError(f"Ticket {ticket_id} not found")
Atomic groups: transactions¶
Multiple operations that must succeed or fail together go inside an explicit transaction:
async with backend.transaction():
await backend.execute("INSERT INTO tickets VALUES (...)", ...)
await backend.execute("UPDATE counters SET ...", ...)
# Both committed atomically. Either raised -- both rolled back.
Nested transactions create savepoints. Inner failures roll back to the savepoint without affecting the outer transaction:
async with backend.transaction(): # outer
await backend.execute(...)
try:
async with backend.transaction(): # inner (SAVEPOINT)
await backend.execute(...)
raise SomeError()
except SomeError:
pass # inner rolled back to savepoint, outer continues
await backend.execute(...) # outer commits cleanly
A transaction belongs to the task that opened it. Another command running
at the same time does not join it: that command's writes are not rolled
back with it, and its reads do not see rows the transaction has not
committed. SQLite has one connection, so there the other command waits
for the transaction to end. A task started inside the body shares the
transaction instead of getting its own, so run its statements in the
body or start it after the block ends. On SQLite, such a task opening a
transaction of its own raises RuntimeError.
Only the raw-SQL methods (execute, fetch, executemany,
fetch_one) participate in the transaction. To group namespace
operations (row_upsert, kv_*, etc.) atomically, use raw SQL inside
the transaction body. A namespace write inside the body commits on its
own on PostgreSQL and raises RuntimeError on SQLite, where it would
wait on the lock the transaction holds. A namespace read inside the body
sees the transaction's uncommitted writes on SQLite, which has one
connection, and not on PostgreSQL, where it runs on a connection of its own.
The transaction holds an underlying connection for the lifetime of the
async with block. Long-running transactions starve the pool, and on
SQLite they hold up every other command's database calls. Keep
transaction bodies short.
Portability vs vendor-specific code¶
CascadeUI does not translate or rewrite SQL. Code targeting a specific backend uses that backend's dialect directly:
# PostgreSQL-specific (will not work on SQLite)
await backend.execute(
"CREATE INDEX CONCURRENTLY ix_tickets_user ON tickets(user_id)"
)
For portability across backends, use the documented subset:
- Types:
INTEGER/BIGINT,TEXT,BLOB/BYTEA(different names, same semantic),REAL/DOUBLE PRECISION. AvoidJSONB,TIMESTAMPTZ,ARRAY,ENUM(PostgreSQL-only). - Functions:
COALESCE,LOWER,UPPER,COUNT,MAX,MIN,SUM,AVGare portable. Date/time functions diverge -- store epoch integers and convert at the application layer. - Operators:
=,<>,<,>,<=,>=,LIKE,IS NULLare portable. Vendor operators (@>,?JSONB containment in PostgreSQL;GLOBin SQLite) are not.
When raw SQL is wrong¶
Reach for raw SQL when the namespace API genuinely cannot express what
you need. Anything CascadeUI's UI state covers (cascadeui_persistent_views,
cascadeui_application_slots, cascadeui_kv) should flow through the
namespace API. The escape hatch is for code outside that domain.
InMemoryBackend does not declare Capability.RAW_SQL; tests against
in-memory storage cannot use the raw-SQL surface. Code paths that
require raw SQL skip in-memory testing or use a real-DB fixture
(testcontainers for PostgreSQL, a temp file for SQLite).
Writing a custom backend¶
A backend is any class that satisfies the PersistenceBackend Protocol and
declares its capabilities. No inheritance is required; the Protocol is
@runtime_checkable:
Declare how schema migrations reach your backend
Library releases occasionally migrate the tables they own (the
cascadeui_persistent_views v2 migration adds a nullable column). A migrator needs
one of two declarations from the backend. Capability.OPEN_ROWS says rows
are open mappings: row_upsert accepts unknown columns, row_select
returns them untouched, and reading a missing one gives None -- a
column-add alters nothing on such a backend, so the migrator skips its
DDL. Capability.RAW_SQL with an implemented execute lets the migrator
alter the table instead. A backend declaring neither is rejected with
PersistenceConfigError during PersistenceMiddleware.initialize when a
migration is pending, with both remedies named in the message. A backend
with nothing to migrate is never asked for either flag.
The fresh-install case above is literal: migration only runs on a table
an older release created. A
backend that builds its own fixed-column tables owns keeping that DDL at
the current column set, because a table it creates today is stamped at
today's version and no migrator ever inspects it. Miss a column and the
first registry write fails at the driver, past the reach of the check
above. Building the table from the shared DDL constants in
cascadeui.persistence.schema is the way not to drift.
from cascadeui.persistence import Capability, PersistenceBackend
class MyBackend:
# OPEN_ROWS: rows here are open mappings, so schema migrations skip
# their DDL. A fixed-column backend declares RAW_SQL instead.
capabilities = (
Capability.KV
| Capability.RELATIONAL
| Capability.SCHEMA_META
| Capability.OPEN_ROWS
)
async def initialize(self) -> None: ...
async def close(self) -> None: ...
# Key-value surface (Capability.KV)
async def kv_read(self, namespace, key): ...
async def kv_write(self, namespace, key, value): ...
async def kv_delete(self, namespace, key): ...
async def kv_scan(self, namespace, prefix=""): ...
# Relational surface (Capability.RELATIONAL)
async def row_upsert(self, namespace, row, key_columns): ...
async def row_select(self, namespace, where=None): ...
async def row_delete(self, namespace, where): ...
async def row_delete_where_lt(self, namespace, column, value): ...
# Optional. Omit it and the flush falls back to a row_upsert loop;
# implement it to collapse a flush into one round-trip.
async def row_upsert_many(self, namespace, rows, key_columns): ...
# Schema metadata (Capability.SCHEMA_META)
async def get_schema_version(self, table): ...
async def set_schema_version(self, table, version): ...
Capability flags¶
Each namespace config declares which capabilities it needs; the manager
validates those against the backend's declared set when the middleware
initializes. A mismatch raises PersistenceConfigError before any
backend method runs:
| Namespace | Required capabilities |
|---|---|
RegistryPersistence |
RELATIONAL \| SCHEMA_META |
ApplicationPersistence (no TTL slots) |
RELATIONAL \| SCHEMA_META |
ApplicationPersistence (any ttl_days slot) |
RELATIONAL \| SCHEMA_META \| TTL_INDEX |
OPEN_ROWS and RAW_SQL sit outside the required sets above. They are
checked at one conditional seam instead: when apply_migrations finds a
pending schema migration for a namespace, that namespace's backend must
declare one of the two, and a backend declaring neither raises
PersistenceConfigError at that point. A backend with no migrations
pending is never asked for either.
Declare capabilities on the class, not the instance:
Backend contracts beyond method signatures
Three correctness properties are required of every backend:
- Copy on store:
row_upsertmust not retain a reference to the caller's dict. A later mutation on the caller side must not bleed into storage. - NULL-safe TTL prune:
row_delete_where_ltmust not sweep rows whose target column isNULL. SQL'sNULL < valueevaluates toNULL(never true); the Protocol requires that same semantic from every implementation. - Scan snapshot safety:
kv_scanmust not raiseRuntimeErrorwhen the caller writes to the namespace mid-iteration. Snapshot keys up front.
InMemoryBackend is the reference for all three. Tests in
tests/test_backends.py parametrize the full Protocol surface across
every shipped backend. Drop a custom class into that fixture to see
the same coverage applied to yours.
Slot policies¶
SlotPolicy carries per-slot policy for application slots: opt-in
persistence and an optional TTL. Slots default to ephemeral -- the
policy's persistent=True flag is the opt-in that the persistence
middleware watches for.
from cascadeui.persistence import SlotPolicy
SlotPolicy() # ephemeral (default)
SlotPolicy(persistent=True) # durable, no TTL
SlotPolicy(persistent=True, ttl_days=7) # durable, prune after 7 days
SlotPolicy(ttl_days=7) # ValueError -- TTL needs persistent=True
Declare slot policies in two places:
# (1) Static: inside ApplicationPersistence.slots
application=ApplicationPersistence(
backend=SQLiteBackend("app.db"),
slots={
"preferences": SlotPolicy(persistent=True),
"cache:search": SlotPolicy(persistent=True, ttl_days=7),
},
)
# (2) Runtime: after PersistenceMiddleware has initialized, via the manager
manager = store.persistence_manager
manager.register_slot_policy(
"cache:autocomplete",
SlotPolicy(persistent=True, ttl_days=1),
)
A slot with no registered policy takes SlotPolicy()'s retention, so it
never expires; whether it is saved still follows the three routes below. The
fallback is logged at DEBUG, once per slot.
Three ways to opt a slot in
All three routes mark the slot persistent; pick the one that lives closest to where the slot is defined:
persistent_slots = ("name",)on the view class -- declarative and the recommended default. Registered at class-definition time.SlotPolicy(persistent=True)inApplicationPersistence.slots-- use when persistence is pure config (TTL tuning, no owning view).access_slot(..., persistent=True)from code -- use when the declaration lives next to the slot's seed logic inside a reducer or aseed_initial_statehook.
All three register the slot name in a sticky module-level set, so every later write with the same name inherits the contract.
Choosing a persistence pattern¶
Two axes decide the pattern: whether the data must survive restart, and whether the view must re-attach to its original message.
| Data survives restart? | View re-attaches? | Pattern | Stable persistence_key required? |
|---|---|---|---|
| No | No | Plain StatefulView (no persistence) |
No |
| Yes | No | Pattern 1 (named slot) | Yes -- pass persistence_key= explicitly |
| No | Yes | PersistentView subclass (registry only) |
Yes (registry identity) |
| Yes | Yes | PersistentView plus Pattern 1 slot |
Yes (shared across both roles) |
persistence_key is opt-in identity. The property falls back to self.id
(a fresh UUID per instance) when no persistence_key= is passed at
construction. That fallback is safe for the top row and for views
that never key a persistent slot off self.persistence_key. The three rows
that involve persistence need a domain-stable value (guild id,
composite user-guild key, or an explicit persistence_key=f"counter:{uid}")
to avoid writing to a fresh bucket on every restart.
Pattern 1: Data persistence via a named slot¶
Persistence in this pattern comes from the persistent_slots class
attribute (or access_slot(..., persistent=True)) -- that flag is the
opt-in that tells the library to write the slot to disk. Setting
persistence_key alone persists nothing; it only names the lookup bucket
inside the slot.
The view instance is recreated on each invocation; only the data
survives. The canonical opt-in is declarative: list the slot name in
the view's persistent_slots class attribute, write to the slot from
a reducer, and read back through slot_property.
from cascadeui import (
StatefulLayoutView, cascade_reducer, slot_property, access_slot,
)
@cascade_reducer("COUNTER_INCREMENT")
async def increment(action, state):
slot = access_slot(state, "counters", action["payload"]["key"])
slot["value"] = slot.get("value", 0) + 1
return state
class CounterView(StatefulLayoutView):
instance_limit = 1
persistent_slots = ("counters",)
value = slot_property(
"value", slot="counters", key=lambda self: self.persistence_key, default=0,
)
from cascadeui import (
StatefulView, cascade_reducer, slot_property, access_slot,
)
@cascade_reducer("COUNTER_INCREMENT")
async def increment(action, state):
slot = access_slot(state, "counters", action["payload"]["key"])
slot["value"] = slot.get("value", 0) + 1
return state
class CounterView(StatefulView):
instance_limit = 1
persistent_slots = ("counters",)
value = slot_property(
"value", slot="counters", key=lambda self: self.persistence_key, default=0,
)
Pattern 1 needs a stable persistence_key
slot_property(..., key=lambda self: self.persistence_key) reads the
slot bucket whose name is self.persistence_key. Pass persistence_key=...
explicitly at construction (e.g. persistence_key=f"counter:{user_id}")
or point the key= lambda at a different stable identifier --
guild id, composite user-guild key, or any domain value that
survives reconstruction.
The UUID fallback on persistence_key is fresh per instance. Combining
it with persistent=True writes to a new bucket every restart,
leaving the prior data orphaned on disk.
persistent_slots vs manual opt-in
persistent_slots is shorthand for access_slot(name, persistent=True)
without the seed hook. Every route registers the name in the same
sticky module-level set, so you do not need to re-pass the kwarg
from reducers or other call sites. Use the class attribute as the
default; reach for access_slot(..., persistent=True) only when the
opt-in genuinely belongs next to seed logic (dynamic slot names,
per-invocation decisions).
Pattern 2: View persistence via PersistentView¶
PersistentView (V1) and PersistentLayoutView (V2) stay interactive
across bot restarts:
import discord
from discord.ui import ActionRow
from cascadeui import PersistentLayoutView, StatefulButton, card
class RoleSelectorPanel(PersistentLayoutView):
instance_limit = 1
instance_scope = "guild"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.add_item(
card(
"## Role Selector",
ActionRow(
StatefulButton(
label="Get Role",
custom_id="roles:get",
callback=self.give_role,
),
),
color=discord.Color.blurple(),
)
)
async def give_role(self, interaction): ...
async def on_restore(self, bot): ...
from cascadeui import PersistentView, StatefulButton
class RoleSelectorView(PersistentView):
instance_limit = 1
instance_scope = "guild"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.add_item(StatefulButton(
label="Get Role",
custom_id="roles:get",
callback=self.give_role,
))
async def give_role(self, interaction): ...
async def on_restore(self, bot): ...
Send once from an admin command:
@bot.hybrid_command()
async def setup_roles(ctx):
view = RoleSelectorPanel(context=ctx, persistence_key=f"roles:panel:{ctx.guild.id}")
await view.send()
After a restart, setup_middleware(PersistenceMiddleware(bot=self, ...))
drives the reattach pipeline during startup:
- Reads the registry via
RegistryPersistence.backend.row_select()and records every row instate["persistent_views"], soexit()on a restored panel deletes its row and a re-send under its key retires it, the same as for a panel sent in this process. - Looks up each row's
view_classin the class registry (see Class identity). - Walks the kwargs migrator chain from the stored
kwargs_schema_versionto the class's current version. - Fetches the target channel and message (skips non-messageable channels).
- Constructs the view, sets
_message, restoresuser_id/guild_id/session_idfrom the row, and callsbot.add_view(view, message_id=...). - Registers the view in state, installs the message-deletion listener
(eagerly, since restored views skip
send()), and callson_bind(bot)so runtime dependencies are injected. Interaction routing is live from this point; the view is clickable. - After
setup_hookreturns and the gateway connects, a background task awaitsbot.wait_until_ready()and then callson_restore(bot)on each restored view. The gateway cache is warm at that point, so renders resolve real users, members, and channels instead of cold defaults. Panels sharing a channel render one at a time (message edits rate-bucket per channel); panels in distinct channels render concurrently.
Retiring a registration¶
exit() on the instance that owns the registration removes the panel's
registry row along with the usual teardown. A panel superseded under its key
(see Replacing a panel under its own key)
no longer owns the row, so its exit() leaves it alone. When no instance is live (the message was deleted while the
bot was offline, or the panel is being decommissioned from an admin task),
retire the row directly through the manager:
A panel that navigates keeps its registration through the view now on its
message. Closing that view with its Exit button, a replace() from it, or the
library's own cleanup (an instance-limit replacement, a parent's exit) closes
the panel: the row goes whether the message is frozen or deleted, as it would
for the panel's own exit(). Sending that view again to another message moves
the row with it instead, so Back rebuilds the panel there and a restart
restores it there. Give a pushed screen a Back button when the panel should
stay up behind it.
A pushed view left idle does not time out on the panel's message. It returns to
the panel, which never times out, so the panel keeps answering clicks instead
of freezing until the next restart. A pushed view that defines its own
on_timeout() times out as that says, and one whose return fails (its panel's
class is no longer registered, or the edit is refused) times out as usual,
leaving the panel restorable after a restart.
The two routes are not interchangeable. exit() retires the registration only
while the view still owns it; prune_registry matches by key and deletes
whatever holds it, so it is for a key no live panel holds. Pruning a key a
panel is running under leaves that panel on screen and unrestorable after the
next restart, which the library logs as a warning. Check first with
get_active_view(persistence_key=...), and when both are involved, prune after
the exit rather than before: an exit can hand the registration back to a live
predecessor, and a prune that ran first would delete the row it just restored.
See Pruning for the full prune surface, including the
age-verified prune_unreachable sweep.
Class identity: rows resolve by the name they recorded¶
Step 2 resolves each row by string. The view_class column holds the
qualified class name recorded when the panel was posted
(f"{cls.__module__}.{cls.__qualname__}" by default), and the class registry
is keyed the same way at import. Moving the class to another module, renaming
it, or renaming a parent package changes the key on the import side only.
Existing rows still hold the old string, so the lookup misses: they land in
the skipped bucket, and each posted panel stays in its channel with frozen
content and dead controls.
Nothing breaks at refactor time or in the running process. The failure waits
for the next restart, where the signal is the reattach summary and, once a
reattach() pass has run, a warning naming the unresolved strings. Pin the
stored name before shipping the refactor:
class TicketPanel(PersistentLayoutView): # moved from mybot.views
session_class_key = "mybot.views.TicketPanel" # the name the rows hold
session_class_key replaces the derived name in session IDs, the instance
index, session origin tracking, and the view_class column rows record, so
existing rows keep resolving and new rows record the pin. Navigation entries
are the one string it does not touch: they record the class's import path,
which is what pop() reconstructs by.
- The pin must equal the
view_classvalue the rows already hold: the class'smodule.QualNameat the time they were written. - One pin per class. Two persistent classes sharing a pin raise at class definition, because their rows could not say which class wrote them. If the other class is the old definition of the one you moved, delete it rather than changing the pin.
- The pin does not inherit. A subclass records its own name unless it sets its own pin.
- Kwargs migrators key on the same stored string, so after pinning, register them under the pin rather than the class's new path (see Kwargs migrations).
Rows that already missed a pass are not lost. They stay on disk in the
skipped bucket and recover on the next restart, or immediately via
reattach(), once a class
registers under the stored name.
View identity: user_id and session_id follow the construction context¶
What identity a persistent view carries, and whether it has a session, is decided by the context it is constructed with:
- An interaction or command context (
context=ctx) derives the invoker'suser_id. Auser_idkeys a session: the navigation context that holds the push/pop chain, itsshared_data, and the undo timeline. This is the shape for a per-user persistent panel re-attached to its owner. - A bare channel (
context=channel) derives nouser_id-- adiscord.TextChannelhas no.author. The view is an ownerless guild artifact with no session. This is the shape for a public board anyone in the channel uses: a leaderboard, a role panel, a status display.
Both round-trip through restore intact: the registry row stores user_id when there
is one, and reattach restores it alongside the session_id the row carries (step 5 above).
A channel-posted panel restores ownerless by design -- nothing to attach, nothing to
key a session on.
Restored session IDs come back verbatim
The registry row stores the session_id the view was registered under, and
reattach restores that value as-is, UUID suffix and all. A view that was
MyPanel:user_123:a1b2c3d4 before the restart is MyPanel:user_123:a1b2c3d4
after it, so a stored id still matches the restored view's.
Only a row written before the session_id column existed comes back with
NULL there. Reattach falls back to deriving MyPanel:user_123 from the
restored user_id for those, which is the coalesced shape without a suffix.
owner_only and user_id are independent axes: owner_only governs who may
interact (a public board sets owner_only = False), while user_id governs
identity and session. A public board can still record its posting admin -- pass
user_id=ctx.author.id explicitly alongside the channel; the explicit value is kept,
because derivation only fills a None user_id. Leaving it ownerless is usually the
honest model: the panel belongs to the channel, not a person.
Runtime dependencies via on_bind¶
A persistent view often needs runtime handles (a database pool, the bot itself, a service client) to load its data. These cannot ride the constructor: the registry row is JSON, and a pool or bot is not serializable. Passing one as a constructor kwarg declines the registry write (with a directed error naming the kwarg and pointing here), so the view would silently drop on the next restart.
Dict keys inside a kwarg must be strings for the same reason. JSON stores
every key as a string, so scores={42: 3} restores as {"42": 3} and a
lookup by 42 misses. The row is still written, and an ERROR names the
kwarg and the key.
Inject them in on_bind(bot) instead:
class LeaderboardPanel(PersistentLeaderboardLayoutView):
async def on_bind(self, bot):
await super().on_bind(bot)
self.db = bot.db
async def on_load(self):
# self.db is set -- on_bind ran first
self.entries = await self.db.fetch_standings()
Call super().on_bind(bot) first, since a pattern class binds its own
dependencies there. The bot itself is stored on the view before the hook
runs, and the leaderboard exposes it as self.bot, so an override has no
need to assign it.
The library calls on_bind(bot) automatically at two points: during send()
(when the bot is derivable from the construction context) and during restore,
before on_restore. A view posted with a bare channel context carries no
.bot, so send() cannot derive it. That view calls on_bind itself before
send():
view = LeaderboardPanel(context=channel, persistence_key=f"board:{board_id}")
await view.on_bind(bot) # channel context: no derivable bot
await view.send()
Without that call the hook does not run, and its own dependencies stay
unbound. A PersistentLeaderboardLayoutView still resolves avatars: with no
bot bound, its bot reads the one PersistenceMiddleware holds.
Keep on_bind idempotent; it may run more than once. A sync override
(def on_bind) is also accepted.
Set attributes, not UI side effects
on_bind runs before the view is displayed -- ahead of the first
render at send() time and ahead of on_restore at restore time. A
refresh() or send() from inside on_bind is premature: at send time
there is no message to edit yet, and at restore time it ships an edit
before on_restore has rebuilt the view. Assign the dependencies in this
hook; do the data load and render in on_load or on_restore.
Kwargs migrations for PersistentView subclasses¶
When a PersistentView subclass changes its __init__ signature, bump
kwargs_schema_version on the class and register a migrator. Versions start
at 1, and any value that is not a positive int raises ValueError when the
class is defined:
from cascadeui.persistence import register_kwargs_migrator
class TicketPanel(PersistentLayoutView):
kwargs_schema_version = 2 # was 1 before the rename
...
@register_kwargs_migrator("mybot.views.TicketPanel", from_version=1)
async def migrate_ticket_panel_1_to_2(kwargs):
kwargs["channel_id"] = kwargs.pop("target_channel_id")
return kwargs
The qualified class name must match the stored view_class column: the
class's f"{module}.{cls.__qualname__}" at post time, or its
session_class_key
pin. A migrator registered under a name the rows do not hold never runs, so
after a class move, key it by the pinned name rather than the class's new
path. Rows whose stored version is behind the class with no migrator for the
next step are skipped with a WARNING and left on disk for later recovery.
Stale entry handling:
| Scenario | Outcome |
|---|---|
Message or channel returns a definitive 404 (discord.NotFound) |
Row removed, reattach summary logs as removed |
Message or channel transiently unreachable (Forbidden, RateLimited, HTTPException, a transport failure, or non-messageable) |
Row kept, reattach summary logs as unreachable |
View class not imported, or no class registers under the stored name (a moved or renamed class with no session_class_key pin) |
Row kept, reattach summary logs as skipped |
| Kwargs migrator raises or returns non-dict | Row kept, reattach summary logs as failed |
| Construction raises during reattach | Row kept, reattach summary logs as failed |
on_restore raises (post-ready render) |
View stays registered; failure is logged, not in the summary |
Requirements for PersistentView:
persistence_keyis required and must be a non-empty string: a missing or empty key raisesValueErrorand a non-string oneTypeError, since stored keys come back from the database as strings.- All components must have explicit
custom_idvalues (auto-generated IDs do not survive restarts).send()raisesValueErrorfor one without, whether it was built in__init__or inon_load(). timeoutis forced toNone(persistent views never time out).owner_onlydefaults toFalse; override explicitly if your panel should be creator-only.- Ephemeral sends are rejected (
PersistentView.send(ephemeral=True)raisesValueError; ephemeral messages have no permanent ID).
One message per persistence_key
The registry tracks one message per persistence_key. Sending a second
view with the same key exits the previous instance and overwrites
the row. Design keys to be unique per intended panel instance (for
example, "roles:main" for a single shared panel,
f"profile:{user_id}" for a per-user panel).
Replacing a panel under its own key¶
Re-sending a panel under a key another panel holds retires the old one during
the send: a live instance is exited (its exit_policy decides whether the
message is frozen or deleted), and so is a view it pushed onto that message; a
message left from before a restart is frozen or deleted by the new panel's
exit_policy the same way. That suits a plain re-post.
It does not suit a swap that has to confirm the new panel before the old one
comes down, because by the time send() returns the old panel is already gone.
Set retire_previous_on_send = False to retire the predecessor yourself:
from cascadeui import PersistentLayoutView, get_store
class CardPanel(PersistentLayoutView):
retire_previous_on_send = False
async def repost(key, context):
old_panel = get_store().get_active_view(persistence_key=key)
new_panel = CardPanel(context=context, persistence_key=key)
await new_panel.send()
if not await confirm_new_panel(new_panel):
await new_panel.exit(delete_message=True) # the old panel is untouched
return
if old_panel is not None:
await old_panel.exit(delete_message=True)
get_store().get_active_view(persistence_key=...) finds the live panel for a
key, so a caller never has to keep its own registry of panel instances. Read it
before the send: once the new panel registers, the lookup returns the new one.
The new panel owns the registration as soon as it is sent, so the old panel's
exit() leaves the new row in place, and so does deleting the old message
(which exits the old view through the message-deletion cleanup). When the new
panel exits instead, as in the rollback above, the registration moves back to
the old panel while it is still live, so a restart reattaches the panel that is
actually on screen. That holds when the old panel is showing a view it pushed:
the registration goes back to its message with the panel's own class and
arguments, the ones Back would rebuild it from. A predecessor you have already
stopped is treated as retired, and the new panel's exit() removes the row.
The flag is per class; set_class_attribute("retire_previous_on_send", False)
sets it for one instance.
Retiring a panel this way never involves prune_registry. That call matches by
key and would delete whichever row the key points at, including the one an
exit() just handed back to a live predecessor, so a prune belongs after every
exit and only for a key no live panel holds.
Pattern 3: Click routing via DynamicPersistentButton¶
Some persistent buttons do not need a view at all. A role self-assign
button carries its intent in its custom_id: click handling depends
only on the embedded role ID, not on any session state or view
lifecycle. DynamicPersistentButton is the primitive for that shape.
import discord
from cascadeui import DynamicPersistentButton
class RoleToggleButton(
DynamicPersistentButton,
template=r"roles:(?P<category>[a-z_]+):(?P<role_id>[0-9]+)",
):
def __init__(self, *, category: str, role_id: int):
button = discord.ui.Button(
label=f"Toggle {category}",
custom_id=f"roles:{category}:{role_id}",
style=discord.ButtonStyle.primary,
)
super().__init__(button)
self.category = category
self.role_id = role_id
async def on_click(self, interaction):
member = interaction.user
role = interaction.guild.get_role(self.role_id)
if role in member.roles:
await member.remove_roles(role)
else:
await member.add_roles(role)
Subclasses auto-register at class-definition time. The reattach pass
inside await setup_middleware(PersistenceMiddleware(..., bot=bot))
calls bot.add_dynamic_items(*subclasses) so every
DynamicPersistentButton routes correctly after a restart. No separate
wiring step. A subclass imported after that pass (a cog loaded later)
is wired in by the same reattach() re-drive that recovers
late-imported view classes.
Named capture groups in the template are passed as keyword arguments to
__init__ by the default from_custom_id. Captures named user_id,
guild_id, channel_id, role_id, or message_id auto-coerce to
int; other captures pass through as strings. Override
from_custom_id when the subclass needs custom extraction (non-
snowflake coercion, combined keys, lookup-based restoration).
Pattern 2 vs Pattern 3: which to reach for¶
| If the click... | Reach for | Because |
|---|---|---|
Depends only on IDs encoded in the custom_id |
DynamicPersistentButton |
No view means no memory overhead per button and no lifecycle to manage |
| Needs to read or update Redux state | PersistentView |
Full _StatefulMixin machinery is available; state subscription and refresh() are free |
| Coordinates with other components on the same message | PersistentView |
Components inside a view share access to the view's state |
| Is one of N instances that differ only by an embedded ID | DynamicPersistentButton |
One class + one regex routes every click; no per-instance tracking |
| Needs a timeout or exit lifecycle | PersistentView |
DynamicPersistentButton has no lifecycle -- clicks route forever once registered |
The two patterns compose: a PersistentLayoutView can host
DynamicPersistentButton instances in its ActionRows. Cardinality-
driven patterns like role-assign panels use exactly this shape -- the
view owns layout and category organization, while each role button is
a DynamicPersistentButton so buttons differ only by their encoded
(category, role_id) pair.
Migrations¶
Two migrator surfaces exist:
- Schema migrators (library-owned,
register_migrator): rewrite a backend table from version N to N+1. The registry exists so schema changes have a clean landing spot without another breaking release. - Kwargs migrators (user-owned,
register_kwargs_migrator): rewrite a singlePersistentView's storedinit_kwargsblob from version N to N+1. Pure function of the kwargs dict, no backend access.
from cascadeui.persistence import register_migrator
@register_migrator("cascadeui_persistent_views", from_version=2)
async def _migrate_persistent_views_2_to_3(backend):
rows = await backend.row_select("cascadeui_persistent_views")
for row in rows:
row["new_column"] = derive(row)
await backend.row_upsert("cascadeui_persistent_views", row, ["persistence_key"])
The row API resolves a namespace to a table itself, so the example above
works under any table_prefix. Raw SQL does not. A migrator that reaches
for backend.execute or backend.fetch names its own table, and naming
the logical one there targets the unprefixed table: absent on a prefixed
database, or a consumer's own table of that name. Resolve it first:
from cascadeui.persistence import Capability, physical_table, register_migrator
@register_migrator("cascadeui_persistent_views", from_version=3)
async def _migrate_persistent_views_3_to_4(backend):
if Capability.OPEN_ROWS in backend.capabilities:
return # open rows already read a missing column as None
table = physical_table(backend, "cascadeui_persistent_views")
async with backend.transaction():
await backend.execute(f"ALTER TABLE {table} ADD COLUMN priority INTEGER")
A migrator that changes columns should return early on a backend declaring
OPEN_ROWS, whose rows are open mappings that a new column does not alter.
A backend without RAW_SQL is not that case: it only lacks the raw-SQL
surface. Keep several statements in one transaction(). The new version is
recorded after the migrator returns, so one that fails partway runs again
on the next boot, and so does one interrupted between its commit and that
record. SQLite has no ADD COLUMN IF NOT EXISTS, so a migrator adding a
column there reads it first and returns when it is already present.
Library-owned migrators run automatically during apply_migrations in the
setup pipeline. A missing migrator for a required version step raises
PersistenceSchemaError. Fresh installs skip this path entirely because
the DDL creates tables at the current version; a table that has rows but no
recorded version predates schema versioning and is treated as version 1, so
its migrators still run. A pending migration also
checks the backend's capability declaration: OPEN_ROWS skips the DDL,
RAW_SQL runs it, and neither raises PersistenceConfigError (see
Writing a custom backend).
Registering a migrator under one of the pre-rename table names raises
ValueError naming the current name. persistent_views and
application_slots are not keys the migration loop resolves, so a
migrator keyed on either would sit in the registry and never be looked up.
The same refusal covers the migrators= bulk form, which registers through
the same function.
The library ships one schema migrator: cascadeui_persistent_views
from version 1 to 2, adding the first_unreachable_at column described in
Rows that stay unreachable. A database
created before that column existed migrates automatically on the next
startup against a library new enough to know about it; a database already
migrated to version 2 raises PersistenceSchemaError against an older
library that only knows version 1, so a downgrade is not silent.
Registering migrators as data¶
The decorators above register at import time, which is the right shape when each migrator is a function you wrote. When they are built as data (generated from a table, loaded from a manifest), pass them to the middleware instead:
await setup_middleware(
PersistenceMiddleware(
backend=SQLiteBackend("state.db"),
bot=bot,
migrators={
"schema": {("cascadeui_persistent_views", 2): _migrate_views_2_to_3},
"kwargs": {("mybot.cogs.panel.TicketPanel", 1): _migrate_panel_1_to_2},
},
)
)
Both keys are optional; each maps (name, from_version) to an async callable.
Registration runs before the migrations and rehydration that consume it, and
re-constructing the middleware never raises a duplicate-key error. A
(name, from_version) already registered keeps the migrator it has, and the
one passed here is ignored with a WARNING; the decorators raise ValueError
for the same collision. The library registers
("cascadeui_persistent_views", 1) itself. The two paths write to the same
registry: use whichever fits how the migrator was authored.
Pruning¶
Three prune methods live on the manager, for a scheduled task or an admin
command; DevTools runs prune_unreachable through /cascadeui unreachable:
manager = store.persistence_manager
# Application slots: delete one slot, OR the slots whose expires_at is more
# than 7 days past. A pruned slot leaves the running bot's state too.
await manager.prune_application(slot="cache:search")
await manager.prune_application(older_than_days=7)
# Registry: delete specific persistent-view rows (or everything).
await manager.prune_registry(persistence_keys=["roles:main", "tickets:panel"])
# Registry: delete rows that have stayed unreachable past a cutoff, after
# re-checking each one against Discord.
await manager.prune_unreachable(older_than_days=30)
prune_unreachable is the safe way to clear the backlog described in
Rows that stay unreachable -- it re-verifies
every candidate before deleting anything, so a row that has recovered is
never mistaken for one that never will.
Each prune dispatches a bookkeeping action (APPLICATION_SLOTS_PRUNED,
REGISTRY_PRUNED) so subscribers and hooks observe the deletion without
inferring it from row counts.
A prune or an expiry leaves undo history alone. On a view with
enable_undo = True, undoing a step taken before the slot was removed brings
back what that step changed (for a dict slot, only the keys it touched), and
the slot is saved again.
Automatic TTL sweeping¶
When any slot declares ttl_days, the manager starts a daily background
sweeper during initialization. Once every 24 hours it deletes the rows whose
expires_at has passed (row_delete_where_lt on
cascadeui_application_slots) and drops the same slots from the running
bot's state, so an expired slot stops being served without a restart. The
cadence is fixed, since TTLs are counted in days.
expires_at is an absolute timestamp written at write-time, not a
duration. A slot is written only when what it stores changes, so its expiry
counts from its last change, not from the bot's last action. It survives bot
restarts: a row written with ttl_days=7 two days before a crash still has
five days left after a restart. rehydrate()
runs one prune pass before reading so rows that expired while the bot was
offline are dropped rather than loaded into memory.
prune_application(older_than_days=...) runs the same deletion on demand.
Observability¶
Register hooks on the manager to observe flush cadence and errors without parsing logs:
manager = store.persistence_manager
def on_flush(namespace, upsert_count, delete_count):
metrics.record(f"persistence.{namespace}.upserts", upsert_count)
metrics.record(f"persistence.{namespace}.deletes", delete_count)
def on_error(namespace, exc):
alerts.fire(f"persistence.{namespace}.error", repr(exc))
manager.register_hook("on_flush", on_flush)
manager.register_hook("on_error", on_error)
Hooks run under the middleware's write lock for the namespace they
describe. Keep them fast and non-blocking. The middleware enters
exponential backoff on flush failure (1s, 2s, 4s, 8s, 16s, capped at 60s)
and logs CRITICAL after MAX_RETRIES consecutive failures. A write still
running after 30 seconds is cancelled and counts as a failure. Rows stay
dirty across retries so no writes are lost.
What gets persisted¶
CascadeUI state is ephemeral by default. Nothing reaches disk until
either (a) a slot is opted in via persistent_slots,
access_slot(..., persistent=True), or SlotPolicy(persistent=True),
or (b) a PersistentView subclass is registered. See
Core Concepts - State Topology for the
full tree.
Scoped state is not persisted by default
dispatch_scoped() is a Redux organization pattern, not a persistence
mechanism. Scoped writes live under state["application"]["scoped"]
(or a named scoped_slot) and are dropped on restart unless the
slot is explicitly opted in. Opt in with persistent_slots = ("scoped",)
on the view class for the default scoped bucket, or
persistent_slots = ("my_named_slot",) paired with
scoped_slot = "my_named_slot" for a named bucket. TTLs live on
SlotPolicy(ttl_days=N) at setup time, never on class attributes.
| State section | Persisted? | Namespace |
|---|---|---|
views, sessions, components, modals |
No (ephemeral runtime) | — |
application.<slot> without opt-in |
No | — |
application.<slot> marked persistent=True |
Yes | application |
application.scoped without opt-in |
No | — |
application.scoped marked persistent=True |
Yes | application |
persistent_views registry |
Yes | registry |
Scoped state (dispatch_scoped) lives under
state["application"]["scoped"], which means it uses the same
persistence plumbing as every other application slot. Opt in once with
persistent_slots = ("scoped",) on a view class (or
access_slot(state, "scoped", ..., persistent=True) inside
seed_initial_state) and scoped writes flow through
ApplicationPersistence to the backend. No mirror-into-a-slot dance
required.
Store IDs, not discord.py objects
State is serialized as JSON. discord.py model objects (Member,
Role, Channel, etc.) are not JSON-serializable: the slot holding
one is not saved, and PersistenceMiddleware logs an ERROR naming it.
Store the .id integer:
# Wrong: the slot is not saved
await self.dispatch_scoped({"target": interaction.user})
# Right: store the snowflake ID
await self.dispatch_scoped({"target_id": interaction.user.id})
An ID used as a dict key goes in as str(user.id). JSON stores every
key as a string, so an int key reads back as "42" after a restart;
the slot is still saved, with an ERROR naming the key.