KPIs β 100 Metrics
Traffic-light against each target. Most are pending this early; values flip as waves land. 2 catalog metrics are now backed by the live worker feed and marked live.
Pending live feed
Worker snapshots not yet readable from this environment β every catalog KPI below is shown as a pending estimate against target.
100
Total
28
On Track
0
At Risk
1
Critical
71
Pending
π΄ MVP01 β The Visible Spine (50)
Pipeline Β· UI Β· validation Β· realtime
Infrastructure & Core
CI/CD build success %
The percentage of successful builds in the continuous integration and delivery pipeline, indicating the reliability and stability of the infrastructure and core systems.
01A1
Avg Coolify deploy time
The average time it takes to deploy Coolify, with lower values indicating better performance and a target of minimizing this duration to ensure efficient and rapid deployment.
01A2
DB tables built vs plan (Set A+B)
The percentage of database tables that have been successfully built against the planned set (A+B), aiming for a 100% completion rate to ensure infrastructure and core systems are fully operational.
01A3
% tables covered by RLS
The percentage of tables in the database that have Row-Level Security (RLS) implemented, aiming to ensure that all sensitive data is properly secured and access-controlled.
01A4
% i4F seed events loading OK
The percentage of i4F seed events that are successfully loaded, with a target of achieving a 100% success rate to ensure reliable infrastructure and core functionality.
01A5
% shadcn/Tailwind UI components done
The percentage of completed shadcn/Tailwind UI components, tracking progress towards the target of 100% completion to ensure a comprehensive and fully-functional infrastructure and core system.
01A6
DEV environment uptime
The percentage of time the development environment is operational and accessible, with a target of exceeding 99% to ensure minimal disruptions to development workflows.
01A7
PROD migration time (supabase db push)
The time it takes to successfully migrate changes to the production Supabase database, with the goal of minimizing this duration to ensure efficient and timely deployment of updates.
01A8
Velocity vs sprint plan
The percentage of planned work actually completed during a sprint, with a target of 100%, indicating the team's ability to deliver on their infrastructure and core commitments.
01A9
% feature-flags implemented
The proportion of feature flags that have been successfully implemented in the infrastructure and core systems, aiming for 100% coverage to ensure efficient and controlled feature management.
01A10
Validation Engine & State
Gates implemented (of 6)
The number of validation gates successfully implemented out of a total of six, tracking progress towards full implementation of the validation engine and state framework.
01B11
Resource-Capacity-Gate accuracy
The percentage of accurate assignments of resource capacity gates, with a target of 100%, to ensure that the validation engine and state are correctly allocating and managing resources.
01B12
Schedule-Conflict-Gate accuracy
The percentage of correctly identified schedule conflicts by the validation engine, with a target of 100% accuracy to ensure seamless and efficient workflow management.
01B13
fn_evaluate_gates latency
The time taken to evaluate gates during the validation process, aiming to ensure it remains below the target threshold of 50 milliseconds.
01B14
% legal transitions approved by DB
The percentage of valid state transitions that are successfully approved by the database, with a target of achieving 100% approval to ensure data consistency and integrity.
01B15
Illegal-transition reject rate
The percentage of invalid state transitions that are correctly rejected by the validation engine, with a target of 100% indicating that all such transitions are properly identified and blocked.
01B16
state_transition_log write speed
The rate at which the system writes state transition logs, with the goal of achieving a fast write speed to ensure efficient and timely recording of state changes.
01B17
UIβserver drift rate
The percentage of discrepancies between the user interface and server states, aiming to ensure data consistency and accuracy across the system.
01B18
Ought-ness Grid compute per node
The computational efficiency of each node in the grid, aiming for a fast processing time to ensure optimal performance of the Validation Engine and State.
01B19
Ceiling enforcement (blocks >62% on CRITICAL)
The effectiveness of the validation engine in enforcing the ceiling rule by tracking the percentage of blocks that exceed 62% on critical paths, with the target of passing the validation check.
01B20
Realtime & Auth
Initial WS connect time
The time it takes for a user to establish an initial WebSocket connection, with lower values indicating better performance and a more responsive user experience.
01C21
Cross-node message latency
The time it takes for a message to be transmitted and processed between different nodes in the system, with a target of less than 300 milliseconds.
01C22
Magic-link email success %
The percentage of successfully delivered magic-link emails, aiming to exceed a target of 98% to ensure reliable and efficient authentication processes.
01C23
6-level Ack response time
The time it takes for the system to respond to a 6-level acknowledgement request, with lower values indicating better performance.
01C24
Error % on isolated vendor/node channels
The percentage of errors occurring on isolated vendor or node channels, with a target of 0% indicating no errors, to ensure reliable and fault-free communication in real-time and authentication processes.
01C25
CRITICAL/WARNING pushed-to-UI in realtime
The percentage of critical and warning notifications that are successfully pushed to the user interface in real-time, with a target of achieving 100% delivery.
01C26
% presence events synced to LIVE gate
The percentage of presence events successfully synced to the LIVE gate within the Realtime & Auth group, aiming for a target of 100% synchronization.
01C27
Realtime packet-loss rate
The percentage of packets lost during real-time communication, with a target of approximately 0% to ensure high-quality, uninterrupted service.
01C28
Reconnection time after forced drop
The average time it takes for a user to successfully reconnect to the system after being disconnected due to a forced drop, with the target being to minimize this time.
01C29
Duplicate-message rate in tasks (idempotency)
The percentage of duplicate messages received in tasks, indicating the effectiveness of idempotency mechanisms in preventing redundant processing.
01C30
Test Coverage
validation-engine unit coverage
The percentage of the validation engine's codebase that has been tested, aiming for a target of >88% to ensure comprehensive testing and reliability.
01D31
UI component coverage
The percentage of user interface components within the application that are covered by automated tests, ensuring that a significant portion of the UI is thoroughly tested.
01D32
% RPC server-fns covered
The percentage of RPC (Remote Procedure Call) server functions that are covered by automated tests, indicating the extent to which the server's functionality is being tested and validated.
01D33
Passing Playwright E2E scenarios
The percentage of End-to-End (E2E) scenarios written using Playwright that pass successfully, indicating the effectiveness of automated testing in ensuring the application's functionality.
01D34
Full suite runtime
The total time it takes to execute the entire test suite, with a target of minimizing this duration to optimize testing efficiency.
01D35
% negative-path scenarios passing
The percentage of test scenarios that successfully execute and pass when simulating error or invalid conditions, aiming to ensure comprehensive coverage of potential failure paths.
01D36
Open Blocker/Critical bugs
The number of open blocker or critical bugs, with a target of 0, indicating that all high-severity issues should be resolved to ensure thorough test coverage and reliable software quality.
01D37
MTTR
Mean Time To Repair (MTTR) measures the average time taken to resolve and repair defects or issues found during testing, with a lower value indicating more efficient defect resolution and better test coverage.
01D38
Regression rate
The percentage of previously passing tests that fail after changes are made to the code, with a lower rate indicating better test coverage and stability.
01D39
TypeScript strictness coverage
The percentage of TypeScript code that is covered by strict type checking, aiming to ensure that all code adheres to the highest level of type safety and maintainability.
01D40
Performance & UX
Home FCP
The "Home FCP" (First Contentful Paint) KPI measures the time it takes for the initial content on the home page to be rendered and visible to the user, aiming to achieve a fast user experience.
01E41
Action-Board TTI
The time it takes for users to take action on a specific task or goal after interacting with an action board, indicating the speed and efficiency of the user experience.
01E42
Initial JS bundle
The size of the initial JavaScript bundle loaded by the browser, aiming to optimize page load times and improve user experience by keeping it under 0KB.
01E43
Framer animation FPS
The Framer animation FPS KPI measures the average frames per second (FPS) of animations within the Framer application, ensuring a smooth and seamless user experience by targeting a minimum of 60 FPS.
01E44
Health-Timeline load (>100 rows)
The time it takes for the system to load a timeline with more than 100 rows, aiming to achieve a fast loading experience.
01E45
Lighthouse a11y
The accessibility of a website or application, aiming to achieve 100% compliance with accessibility standards to ensure equal access for users with disabilities.
01E46
Responsive mobile fit
Responsive mobile fit:** Ensures the website displays and functions correctly on various mobile devices and screen sizes, providing an optimal user experience.
01E47
Console errors
The number of errors reported in the browser console, aiming to achieve a target of 0 to ensure a seamless and(frames) error-free user experience.
01E48
Suggested-Action accuracy per blocker
Suggested-Action accuracy per blocker" measures the percentage of times suggested actions are correctly implemented by blockers, indicating the effectiveness of the system in providing useful recommendations.
01E49
Server Apdex score
The Server Apdex score measures the percentage of user requests that are served within a predetermined threshold (usually 500ms) to evaluate the responsiveness and performance of the server.
01E50
π΅ MVP02 β Operational Depth (50)
Engines Β· notifications Β· security Β· offline Β· DR
Core Engines & Workers
Coverage of all 7 blockers
The percentage of instances where all 7 blockers are covered, indicating the team's ability to address all critical issues within the Core Engines & Workers group.
02A1
Live critical-path recalc speed
The speed at which the system recalculates and updates the critical path in real-time, indicating how quickly it can adapt to changes and ensure optimal performance.
02A2
CASCADES_TO impact accuracy
CASCADES_TO impact accuracy" measures the extent to which predictions made by the Core Engines & Workers are influenced by the cascading effects of their outputs, aiming for high accuracy in downstream processes.
02A3
pg-boss delayed-jobs queue depth
The number of delayed jobs in the PostgreSQL database, indicating the workload pending execution by the pg-boss engine, with a target of being within a healthy range to ensure efficient job processing.
02A4
Escalation-timer firing lag (T+15/45)
The delay between the scheduled escalation time (T+15 or T+45 minutes) and when the actual escalation event is triggered within the Core Engines & Workers group, with the goal of minimizing this lag.
02A5
Background job failure rate
The percentage of background jobs that fail to complete successfully within the Core Engines & Workers group, aiming for a minimal failure rate to ensure system stability and efficiency.
02A6
DLQ fill rate
The DLQ fill rate KPI measures the percentage of messages that end up in the Dead Letter Queue, indicating the success rate of message processing within Core Engines & Workers, with a target of keeping this rate **low** to minimize failures.
02A7
AWAY/OFFLINE sweep speed
The speed at which the system can transition from an active "AWAY" or "OFFLINE" state to a ready operational state, indicating the efficiency of its recovery and activation processes.
02A8
Auto-transitionβNEEDS_WORK accuracy
The accuracy of automatically transitioning tasks to the "NEEDS_WORK" status within the Core Engines & Workers group, with the target being a high rate of correct transitions.
02A9
Vendor-conflict detection accuracy
The accuracy of our systems in correctly identifying and flagging instances where vendor actions or claims conflict with established policies or agreements, aiming for a high rate of true positives and minimal false positives.
02A10
Advanced Notifications
Web Push delivery rate
Web Push delivery rate:** Measures the percentage of web push notifications successfully delivered to users' browsers, aiming for over 95% to ensure broad message reach.
02B11
SMS/WhatsApp delivery latency
The time it takes for SMS and WhatsApp notifications to be delivered to recipients, aiming for a low latency to ensure timely and efficient communication.
02B12
3rd-party API error %
The percentage of notifications that failed due to errors encountered when interacting with third-party APIs, with the target being to keep this percentage as low as possible to ensure reliable notification delivery.
02B13
Quiet-hours enforcement accuracy
Quiet-hours enforcement accuracy" measures the percentage of instances where the system correctly enforces quiet hours, ensuring that notifications are suppressed during designated periods.
02B14
Self-hosted SMTP uptime
The percentage of time that the self-hosted SMTP server is available and functioning correctly, with a target of achieving an uptime of greater than 99%.
02B15
Email-digest generation speed
The time it takes to generate email digests, aiming to achieve a fast processing speed to minimize delays in notification delivery.
02B16
i18n EN/HE/HI correctness
The accuracy of internationalization (i18n) translations for English (EN), Hebrew (HE), and Hindi (HI) within the Advanced Notifications group, aiming for a 100% correctness rate.
02B17
Multi-level recipient-chain success
The percentage of email notifications successfully delivered through complex recipient chains involving multiple levels of recipients, with a target of 100% delivery success.
02B18
Notification-bundling effectiveness
The percentage of notifications that are successfully bundled together, allowing users to view multiple updates in a single notification, with a target of achieving high bundling efficiency.
02B19
VAPID registration time
The VAPID registration time KPI measures the average time it takes for a VAPID (Visible Alternatives to Passwords, Internet-Draft) registration to be completed, aiming to keep this metric low to ensure efficient and timely notification delivery.
02B20
Security & Immutability
% state-changes in event_log/audit_log
The percentage of state changes in the event log/audit log that are successfully recorded, aiming for a target of 100% to ensure complete visibility and accountability of system activities.
02C21
Blocked UPDATE/DELETE on immutable tables
The share of UPDATE/DELETE attempts on the append-only ledgers (audit, override, escalation, event log) that are correctly rejected by both the grant and trigger layers β the floor of tamper-evidence.
02C22
2FA/TOTP success (Producer/Admin)
The percentage of successful two-factor authentication (2FA) using Time-Based One-Time Password (TOTP) attempts made by producers and administrators.
02C23
Session-timeout accuracy
The exactness with which the system enforces session timeouts, ensuring that user sessions are terminated promptly when the specified timeout period has elapsed.
02C24
High/Critical pen-test findings
The number of high or critical security vulnerabilities identified during penetration testing, with a target of zero to ensure the system's security and immutability.
02C25
Optimistic-lock conflicts resolved
The percentage of optimistic-lock conflicts resolved in the system, indicating the effectiveness of the security and immutability measures in place to prevent data inconsistencies.
02C26
% duplicate requests blocked (idempotency)
The percentage of duplicate requests successfully prevented from being processed, ensuring idempotency and protecting against unintended side effects.
02C27
Snapshot-isolation availability
Snapshot-isolation availability** ensures that all read operations within a transaction see a consistent, unchangeable view of the data as it existed at the moment the transaction began, preventing data corruption and race conditions.
02C28
Override-log min-chars enforcement
Ensure that all override logs meet the minimum character requirement to maintain data integrity and prevent trivial or malicious log entries.
02C29
Brute-force blocks
Brute-force blocks** are a security measure that automatically prevents repeated login attempts from a single IP address or user account after a predefined number of failures, thus thwarting unauthorized access.
02C30
Offline Sync & PWA
PWA/service-worker install success
The percentage of users who successfully install the Progressive Web App (PWA) and its associated service worker, indicating a seamless transition to offline functionality and enhanced user experience.
02D31
Local cache hit ratio
The percentage of data requests that are satisfied from the local cache, indicating the effectiveness of the application's ability to leverage offline capabilities and reduce reliance on network requests.
02D32
Offline write-queue size
The number of write operations that are currently waiting to be synchronized with the server in an offline state, aiming to keep this queue size within an acceptable limit to ensure data consistency and responsiveness.
02D33
Reconciliation success on reconnect
The percentage of successful data reconciliations when a user reconnects to the network after being offline, aiming for a target of 100% success rate.
02D34
Critical-first ordering accuracy
The percentage of critical items that are accurately ordered and reflected in the offline system and Progressive Web App, ensuring users have access to the most important product information without delay.
02D35
Read-only-mode switch time on outage
The time it takes to switch to read-only mode during an outage, aiming for a swift transition to maintain data integrity and user experience.
02D36
Data usage under 3G budget
The average data usage per hour under the 3G budget for Offline Sync and PWA features, aiming to maintain a usage rate below a specified threshold in megabytes per hour (MB/h).
02D37
No-data-loss recovery after offline close
The percentage of instances where data is successfully recovered without loss after a device or application goes offline and then reconnects.
02D38
Delta-request compute speed
The time it takes for the application to process and compute delta requests, aiming for a fast performance to ensure seamless offline sync and PWA functionality.
02D39
Sync conflict-resolution errors
The number of synchronization conflicts that could not be automatically resolved, indicating potential data inconsistencies for users operating offline.
02D40
Observability & Load/DR
% systems monitored (GlitchTip/PostHog)
The percentage of systems being monitored through GlitchTip and PostHog, aiming for 100% coverage to ensure comprehensive observability and load/distributed resource (DR) management.
02E41
Prometheus scrape latency
The time it takes for Prometheus to collect and process metrics from targeted sources, indicating the efficiency of the monitoring system.
02E42
Stability at 50 concurrent (k6)
The system's ability to maintain stability under heavy load conditions, specifically when 50 concurrent users are simulated.
02E43
Daily pg_dump backup success
The percentage of days where a successful pg_dump database backup is completed, indicating the reliability of the organization's data backup process.
02E44
Restore RTO drill
The time it takes to recover from a disaster recovery (DR) scenario, with a target of restoring services within less than 5 minutes.
02E45
Actual RPO via WAL PITR
The actual Recovery Point Objective (RPO) achieved through Write-Ahead Logging (WAL) Point-in-Time Recovery (PITR), aiming for a low value to minimize data loss in the event of a failure.
02E46
VPS memory/CPU at LIVE peak
The headroom available on virtual private servers (VPS) during peak live usage, indicating the excess capacity beyond maximum requirements.
02E47
OTel trace accuracy on heavy queries
The percentage of OTel (OpenTelemetry) traces that are accurately traced ornament heavy queries, indicating the reliability of the observability system in capturing critical performance data.
02E48
Supavisor connection-pool load
The health of the connection pool utilized by the Supavisor, indicating whether it is efficiently managing connections and handling load.
02E49
Survival on sudden Supabase restart
The ability of Supabase to continue serving requests and maintaining data integrity after an unexpected restart, ensuring minimal disruption to users.
02E50