Exceptions#
The Exceptions page lists the exceptions CommandCenter and the BenefitManager environments have logged, newest first, a page at a time. The rows appear while they are being read: the API streams them one line at a time instead of answering with one large document.
There are two sources, each switched on or off per environment under Configuration → Features (see Server features):
| Feature | Source | Status |
|---|---|---|
exceptions-sql — Exceptions (database) |
the exception log table TblExceptionLog in the CommandCenter database |
the page at /exceptions |
exceptions-log-analytics — Exceptions (Log Analytics) |
exception entries in Azure Log Analytics (AppTraces) |
the page at /exceptions-log-analytics |
What a page shows#
Each row is one occurrence of an exception: when it happened, the environment (customer and stage) it belongs to, the exception type, the first 500 characters of the message, and its guid. The stack trace and the inner exceptions are loaded when a row is opened.
A row never carries an environment's SQL instance, database, url or single sign-on settings. (The legacy page sent the whole environment record with every row.)
Order. Rows are ordered by the log's row number, which is the order they were written. Exceptions brought in by a customer-data import are written at import time, so they appear at the top as a group even though their dates are older. Whether to order by date instead is an open decision.
Dates. The log stores a date and time without a time zone. The legacy writer stores the IIS
host's local time, so WebApi reads the stored value as Amsterdam time (Exceptions:Sql:StoredTimeZone,
default Europe/Amsterdam; set UTC for an environment whose writer stores UTC). On the wire every
time carries its offset, for example 2026-09-26T14:19:11.09+02:00. In the autumn hour that happens
twice, a stored time is read as the first (summer time).
Using the page#
- Filters sit above the table: the environment (or Without environment), a time range and a search box (type, message or guid; it searches once you stop typing). Clear filters resets them. Clicking an environment name in a row filters on that environment.
- The filters, the page size and the position are part of the address, so a copied link opens the same view and the browser's Back button goes to the previous page.
- The table fills in while the page is read; the line above it says how many rows have arrived and, at the end, how long it took. Changing a filter or the page halfway stops the page being read.
- Open a row (the arrow) for its inner exceptions and stack traces. Copy puts the whole exception on the clipboard; the chat button sends it to the AI assistant's exception analysis.
Paging and the count#
A page is 50, 100, 250 or 500 rows. Older asks for the rows before the last one shown, Newer for the rows after the first one: the position is a row number, not a page number, so the fortieth page is as fast as the first and a new exception arriving meanwhile does not shift the pages.
The total number of matching exceptions is counted after the rows have been sent, with a time
budget of its own (Exceptions:CountTimeout, default five seconds). When the count takes longer,
the page says the count is unavailable; the rows are complete either way. The count also says how
many matching exceptions are newer than the page, which is what "page 3 of 12" is computed from.
Filters#
| Parameter | Meaning |
|---|---|
from, to |
Occurred at or after from, and before to. ISO-8601 instants, e.g. 2026-09-26T08:00:00Z. |
stage |
One environment, by its identifier; none for exceptions without an environment. |
q |
A guid matches that guid exactly. Other text (2 to 200 characters) matches the type or the message; % and _ are taken literally. |
pageSize |
1 to 500 (default 100). |
before / after |
The cursor from the previous page: older than / newer than that row. Never both. |
A text search reads through the messages, so on a large log it is slower than the other filters; it stops as soon as it has a page, and only the count reads everything (within its budget).
The stream#
GET /api/exceptions/sql/stream answers application/x-ndjson: one JSON object per line.
{"type":"page","source":"sql","pageSize":100,"direction":"older"}
{"type":"row","row":{"id":"79633","key":79633,"guid":"818c7eb7-…","occurredAt":"2026-09-26T14:19:11.09+02:00","exceptionType":"System.Reflection.TargetInvocationException","message":"…","messageTruncated":false}}
{"type":"cursor","older":"79534","newer":"79633","hasOlder":true,"hasNewer":false}
{"type":"total","count":18683,"newerCount":0,"timedOut":false,"elapsedMs":24}
{"type":"done","rows":100,"elapsedMs":34}
- Every refusal (a bad date, both cursors, an unknown environment, a switched-off feature) comes before the first line, as an ordinary 400 or 404 problem document.
- A failure after the first line cannot change the status any more: the stream then ends with an
errorline and nodone. A stream withoutdonewas cut short. - Closing the connection stops the query on the server.
GET /api/exceptions/sql/{key} returns one occurrence with its inner exceptions, in order, stack
traces included. It is addressed by the row number the stream gave, not by the guid: the same guid
is reused by separate occurrences, and the legacy page's "all rows with this guid" mixed them up.
How it fits together#
flowchart LR
user([User]) --> spa[SPA: Exceptions page]
spa -->|/api/exceptions/sql/stream<br/>/api/exceptions/sql/key| bff[BFF]
bff -->|/v2/api/exceptions/…| webapi[WebApi]
webapi -->|TblFeature: switched on?| db[(CommandCenter database)]
webapi -->|TblExceptionLog + stage names| db
legacy[Legacy CommandCenter<br/>other repository] -->|writes exceptions| db
sequenceDiagram
autonumber
participant SPA
participant API as WebApi
participant DB as TblExceptionLog
SPA->>API: GET …/stream?pageSize=100&before=79634
API->>API: validate, feature on? (else 400/404, no body)
API-->>SPA: page
API->>DB: TOP 101 … WHERE Key < 79634 ORDER BY Key DESC
loop each row as it is read
API-->>SPA: row
end
API->>DB: any newer rows?
API-->>SPA: cursor
API->>DB: COUNT (own time budget)
API-->>SPA: total (or count unavailable)
API-->>SPA: done
journey
title Finding out why an environment failed
section Narrow down
Open Exceptions: 5: User
Pick the environment: 5: User
Search for the exception type: 4: User
section Read
Rows appear while loading: 5: User
Go to older pages: 4: User
Open a row for the stack trace: 5: User
The Log Analytics page#
The main API's POST /api/NewExceptionLogApi/WriteExceptionLog writes an exception as a structured
warning, which OpenTelemetry sends to Azure Monitor; it lands in the workspace's AppTraces table.
The Log Analytics page reads those entries with the same stream, filters and paging as the database
page. Differences worth knowing:
- What is kept. The writer stores the exception's type, message, stack trace and only the message of the first inner exception, so the detail shows at most two entries and the inner one has no type or stack trace.
- Time range. A Log Analytics query is always bounded: without a From, the page reads the last
30 days (
Exceptions:LogAnalytics:DefaultWindow). - Search matches whole words (terms) of the type or message, the way Log Analytics indexes them; a fragment in the middle of a word is not found. A guid matches the logged guid exactly.
- Arrival. Log Analytics answers a page in one response, so its rows appear together rather than one by one, and new entries take a few minutes to arrive after they are written.
- Newer after Older is always offered on a page you reached with Older (you came from there), to save a second query per page; it is exact on the first page.
- Prerequisites per environment:
LogAnalytics:WorkspaceIdon WebApi (LogAnalytics__WorkspaceId), and the Log Analytics Reader role for WebApi's managed identity on that workspace. Without the setting the page shows "not configured on this host" and its API answers 503. Then switch Exceptions (Log Analytics) on under Configuration → Features.
The main API no longer reads Log Analytics: its GetEntries and getbysearchoptions calls are gone
(they used a skip operator KQL does not have and could not have returned rows). Writing is unchanged.
Performance and indexes#
The default view reads backwards along the primary key and needs no extra index; measured on the development database, a page of 100 rows plus the count takes about 40 ms. Two lookups may need an index on a large production log:
- opening a row finds the rest of its occurrence by guid;
- a page for one rarely-failing environment reads back until it has enough rows.
Measure these on a production-sized database before adding anything:
SELECT TOP 1 CreateDate, SYSUTCDATETIME() AS utc, SYSDATETIME() AS server_local FROM dbo.TblExceptionLog ORDER BY [Key] DESC;
EXEC sp_helpindex 'dbo.TblExceptionLog';
SELECT COUNT_BIG(*) AS total, SUM(CASE WHEN Sequence = 1 THEN 1 ELSE 0 END) AS outer_rows FROM dbo.TblExceptionLog;
If opening a row takes over 200 ms or a page for one environment over a second, add (idempotently,
as a WebApi migration with an empty Down):
CREATE INDEX IX_TblExceptionLog_Guid ON dbo.TblExceptionLog (Guid) INCLUDE (Sequence, CustomerStageKey);
CREATE INDEX IX_TblExceptionLog_Stage_Key ON dbo.TblExceptionLog (CustomerStageKey, [Key] DESC) WHERE Sequence = 1;
The WebApi model declares a unique index on (Guid, CustomerStageKey). Legacy databases do not have
it, and cannot: an occurrence's inner exceptions share its guid and environment. Never create it
there, and do not rely on it.