Skip to content

Exceptions#

The Exceptions page lists the exceptions CommandCenter and the BenefitManager environments have logged, newest first, a page at a time. The rows appear while they are being read: the API streams them one line at a time instead of answering with one large document.

There are two sources, each switched on or off per environment under Configuration → Features (see Server features):

Feature Source Status
exceptions-sql — Exceptions (database) the exception log table TblExceptionLog in the CommandCenter database the page at /exceptions
exceptions-log-analytics — Exceptions (Log Analytics) exception entries in Azure Log Analytics (AppTraces) the page at /exceptions-log-analytics

What a page shows#

Each row is one occurrence of an exception: when it happened, the environment (customer and stage) it belongs to, the exception type, the first 500 characters of the message, and its guid. The stack trace and the inner exceptions are loaded when a row is opened.

A row never carries an environment's SQL instance, database, url or single sign-on settings. (The legacy page sent the whole environment record with every row.)

Order. Rows are ordered by the log's row number, which is the order they were written. Exceptions brought in by a customer-data import are written at import time, so they appear at the top as a group even though their dates are older. Whether to order by date instead is an open decision.

Dates. The log stores a date and time without a time zone. The legacy writer stores the IIS host's local time, so WebApi reads the stored value as Amsterdam time (Exceptions:Sql:StoredTimeZone, default Europe/Amsterdam; set UTC for an environment whose writer stores UTC). On the wire every time carries its offset, for example 2026-09-26T14:19:11.09+02:00. In the autumn hour that happens twice, a stored time is read as the first (summer time).

Using the page#

  • Filters sit above the table: the environment (or Without environment), a time range and a search box (type, message or guid; it searches once you stop typing). Clear filters resets them. Clicking an environment name in a row filters on that environment.
  • The filters, the page size and the position are part of the address, so a copied link opens the same view and the browser's Back button goes to the previous page.
  • The table fills in while the page is read; the line above it says how many rows have arrived and, at the end, how long it took. Changing a filter or the page halfway stops the page being read.
  • Open a row (the arrow) for its inner exceptions and stack traces. Copy puts the whole exception on the clipboard; the chat button sends it to the AI assistant's exception analysis.

Paging and the count#

A page is 50, 100, 250 or 500 rows. Older asks for the rows before the last one shown, Newer for the rows after the first one: the position is a row number, not a page number, so the fortieth page is as fast as the first and a new exception arriving meanwhile does not shift the pages.

The total number of matching exceptions is counted after the rows have been sent, with a time budget of its own (Exceptions:CountTimeout, default five seconds). When the count takes longer, the page says the count is unavailable; the rows are complete either way. The count also says how many matching exceptions are newer than the page, which is what "page 3 of 12" is computed from.

Filters#

Parameter Meaning
from, to Occurred at or after from, and before to. ISO-8601 instants, e.g. 2026-09-26T08:00:00Z.
stage One environment, by its identifier; none for exceptions without an environment.
q A guid matches that guid exactly. Other text (2 to 200 characters) matches the type or the message; % and _ are taken literally.
pageSize 1 to 500 (default 100).
before / after The cursor from the previous page: older than / newer than that row. Never both.

A text search reads through the messages, so on a large log it is slower than the other filters; it stops as soon as it has a page, and only the count reads everything (within its budget).

The stream#

GET /api/exceptions/sql/stream answers application/x-ndjson: one JSON object per line.

{"type":"page","source":"sql","pageSize":100,"direction":"older"}
{"type":"row","row":{"id":"79633","key":79633,"guid":"818c7eb7-…","occurredAt":"2026-09-26T14:19:11.09+02:00","exceptionType":"System.Reflection.TargetInvocationException","message":"…","messageTruncated":false}}
{"type":"cursor","older":"79534","newer":"79633","hasOlder":true,"hasNewer":false}
{"type":"total","count":18683,"newerCount":0,"timedOut":false,"elapsedMs":24}
{"type":"done","rows":100,"elapsedMs":34}
  • Every refusal (a bad date, both cursors, an unknown environment, a switched-off feature) comes before the first line, as an ordinary 400 or 404 problem document.
  • A failure after the first line cannot change the status any more: the stream then ends with an error line and no done. A stream without done was cut short.
  • Closing the connection stops the query on the server.

GET /api/exceptions/sql/{key} returns one occurrence with its inner exceptions, in order, stack traces included. It is addressed by the row number the stream gave, not by the guid: the same guid is reused by separate occurrences, and the legacy page's "all rows with this guid" mixed them up.

How it fits together#

flowchart LR
    user([User]) --> spa[SPA: Exceptions page]
    spa -->|/api/exceptions/sql/stream<br/>/api/exceptions/sql/key| bff[BFF]
    bff -->|/v2/api/exceptions/…| webapi[WebApi]
    webapi -->|TblFeature: switched on?| db[(CommandCenter database)]
    webapi -->|TblExceptionLog + stage names| db
    legacy[Legacy CommandCenter<br/>other repository] -->|writes exceptions| db
Hold "Alt" / "Option" to enable pan & zoom
sequenceDiagram
    autonumber
    participant SPA
    participant API as WebApi
    participant DB as TblExceptionLog
    SPA->>API: GET …/stream?pageSize=100&before=79634
    API->>API: validate, feature on? (else 400/404, no body)
    API-->>SPA: page
    API->>DB: TOP 101 … WHERE Key < 79634 ORDER BY Key DESC
    loop each row as it is read
        API-->>SPA: row
    end
    API->>DB: any newer rows?
    API-->>SPA: cursor
    API->>DB: COUNT (own time budget)
    API-->>SPA: total (or count unavailable)
    API-->>SPA: done
Hold "Alt" / "Option" to enable pan & zoom
journey
    title Finding out why an environment failed
    section Narrow down
      Open Exceptions: 5: User
      Pick the environment: 5: User
      Search for the exception type: 4: User
    section Read
      Rows appear while loading: 5: User
      Go to older pages: 4: User
      Open a row for the stack trace: 5: User
Hold "Alt" / "Option" to enable pan & zoom

The Log Analytics page#

The main API's POST /api/NewExceptionLogApi/WriteExceptionLog writes an exception as a structured warning, which OpenTelemetry sends to Azure Monitor; it lands in the workspace's AppTraces table. The Log Analytics page reads those entries with the same stream, filters and paging as the database page. Differences worth knowing:

  • What is kept. The writer stores the exception's type, message, stack trace and only the message of the first inner exception, so the detail shows at most two entries and the inner one has no type or stack trace.
  • Time range. A Log Analytics query is always bounded: without a From, the page reads the last 30 days (Exceptions:LogAnalytics:DefaultWindow).
  • Search matches whole words (terms) of the type or message, the way Log Analytics indexes them; a fragment in the middle of a word is not found. A guid matches the logged guid exactly.
  • Arrival. Log Analytics answers a page in one response, so its rows appear together rather than one by one, and new entries take a few minutes to arrive after they are written.
  • Newer after Older is always offered on a page you reached with Older (you came from there), to save a second query per page; it is exact on the first page.
  • Prerequisites per environment: LogAnalytics:WorkspaceId on WebApi (LogAnalytics__WorkspaceId), and the Log Analytics Reader role for WebApi's managed identity on that workspace. Without the setting the page shows "not configured on this host" and its API answers 503. Then switch Exceptions (Log Analytics) on under Configuration → Features.

The main API no longer reads Log Analytics: its GetEntries and getbysearchoptions calls are gone (they used a skip operator KQL does not have and could not have returned rows). Writing is unchanged.

Performance and indexes#

The default view reads backwards along the primary key and needs no extra index; measured on the development database, a page of 100 rows plus the count takes about 40 ms. Two lookups may need an index on a large production log:

  • opening a row finds the rest of its occurrence by guid;
  • a page for one rarely-failing environment reads back until it has enough rows.

Measure these on a production-sized database before adding anything:

SELECT TOP 1 CreateDate, SYSUTCDATETIME() AS utc, SYSDATETIME() AS server_local FROM dbo.TblExceptionLog ORDER BY [Key] DESC;
EXEC sp_helpindex 'dbo.TblExceptionLog';
SELECT COUNT_BIG(*) AS total, SUM(CASE WHEN Sequence = 1 THEN 1 ELSE 0 END) AS outer_rows FROM dbo.TblExceptionLog;

If opening a row takes over 200 ms or a page for one environment over a second, add (idempotently, as a WebApi migration with an empty Down):

CREATE INDEX IX_TblExceptionLog_Guid ON dbo.TblExceptionLog (Guid) INCLUDE (Sequence, CustomerStageKey);
CREATE INDEX IX_TblExceptionLog_Stage_Key ON dbo.TblExceptionLog (CustomerStageKey, [Key] DESC) WHERE Sequence = 1;

The WebApi model declares a unique index on (Guid, CustomerStageKey). Legacy databases do not have it, and cannot: an occurrence's inner exceptions share its guid and environment. Never create it there, and do not rely on it.