Health Checks

Health checks let a service report whether it — and the things it depends on (a database, a downstream HTTP API, a queue) — are in a state where it can do its job.

Overview

A service built on Benzene almost always depends on other resources: a database, a storage location, or another service. Health checks let you verify that a service has access to everything it needs and is in the right state to operate.

In Benzene, a health check isn't a special HTTP-only concept bolted onto the framework — it's just another topic in the middleware pipeline. .UseHealthCheck(topic, ...) adds a middleware that intercepts messages for that topic (and the default "healthcheck" topic), runs your registered IHealthChecks, and returns an aggregated IHealthCheckResponse. Because it's ordinary middleware, it works identically across every transport Benzene supports — AWS Lambda (API Gateway, SNS, SQS, Kafka), Azure Functions, ASP.NET Core, or a self-hosted worker — with no special endpoint plumbing beyond the same fluent .Use*() calls you use everywhere else.

Installation

Package What it adds
Benzene.HealthChecks.Core The abstractions: IHealthCheck, IHealthCheckResult, IHealthCheckResponse<T>, HealthCheckStatus, IHealthCheckBuilder, IHealthCheckFactory. Pulled in transitively by the packages below.
Benzene.HealthChecks The processor, HealthCheckBuilder, the .UseHealthCheck(...) middleware, and the built-in SimpleHealthCheck/InlineHealthCheck/FailedHealthCheck checks.
Benzene.HealthChecks.Http HttpPingHealthCheck — pings a downstream HTTP URL.
Benzene.HealthChecks.EntityFramework DatabaseConnectionHealthCheck<TDbContext> and DatabaseHealthCheck<TDbContext> for EF Core.

Add Benzene.HealthChecks to any project that wires up a pipeline; add Benzene.HealthChecks.Http and/or Benzene.HealthChecks.EntityFramework only where you need those specific checks.

Basic usage

using Benzene.HealthChecks;

app.UseHealthCheck("healthcheck", x => x
    .AddHealthCheck<SimpleHealthCheck>()
    .AddHealthCheck(resolver => resolver.GetService<SimpleHealthCheck>())
    .AddHealthCheck("inline", _ => true)
    .AddHealthCheck("inline", async _ => await Task.FromResult(true))
    .AddHealthCheck(_ => true));

Send it a health check request the same way you'd send any other message — for a BenzeneMessage-based transport (SNS/SQS/Kafka), that's a message with the matching topic:

{
  "topic": "healthcheck"
}

The response is a JSON payload describing whether the service is healthy overall and the result of each individual check (see Response format below).

Core concepts

IHealthCheck

public interface IHealthCheck
{
    string Type { get; }
    Task<IHealthCheckResult> ExecuteAsync();
}

Type is the name used to identify the check in the response — it's the value used as (or to derive) the check's key in the response dictionary. Everything else about the interface is intentionally minimal: no context object is passed in, so a check gets whatever it needs (a DbContext, an HttpClient, ...) via constructor injection instead.

IHealthCheckResult / HealthCheckResult

public interface IHealthCheckResult
{
    string Status { get; }
    string Type { get; }
    IDictionary<string, object> Data { get; }
}

Data is a free-form metadata bag — put whatever's useful for diagnosing a failure in there (e.g. "CanConnect", "Url", "StatusCode"). HealthCheckResult implements the interface and exposes static factory helpers instead of a public constructor for the common cases:

HealthCheckResult.CreateInstance(bool success);                                   // Type = "Unknown", empty Data
HealthCheckResult.CreateInstance(bool success, string type);
HealthCheckResult.CreateInstance(bool success, string type, IDictionary<string, object> data);
HealthCheckResult.CreateInstance(Task<bool> success, string type);                // async overload, returns Task<IHealthCheckResult>
HealthCheckResult.CreateWarning(string type);
HealthCheckResult.CreateWarning(string type, IDictionary<string, object> data);

CreateInstance maps success to HealthCheckStatus.Ok or HealthCheckStatus.Failed; CreateWarning always produces HealthCheckStatus.Warning.

HealthCheckStatus

Three string constants, used as the Status value:

Constant Value Effect on IsHealthy
HealthCheckStatus.Ok "ok" Healthy
HealthCheckStatus.Warning "warning" Still counted as healthy
HealthCheckStatus.Failed "failed" Flips the aggregate response to unhealthy

Only Failed affects the aggregate result — a check that reports Warning shows up in the response so you can see it, but doesn't take the whole service down.

IHealthCheckResponse<T> / HealthCheckResponse

public interface IHealthCheckResponse<THealthCheckResult> where THealthCheckResult : IHealthCheckResult
{
    bool IsHealthy { get; }
    IDictionary<string, THealthCheckResult> HealthChecks { get; }
}

This is the payload returned by .UseHealthCheck(...)IsHealthy is true only if every registered check's Status is not Failed; HealthChecks maps a (de-duplicated, see naming below) name to each check's result.

Registering health checks: IHealthCheckBuilder

public interface IHealthCheckBuilder
{
    IHealthCheckBuilder AddHealthCheck<THealthCheck>() where THealthCheck : class, IHealthCheck;
    IHealthCheckBuilder AddHealthCheck(Func<IServiceResolver, IHealthCheck> func);
    IHealthCheck[] GetHealthChecks(IServiceResolver resolver);
}

There are two fundamentally different ways a check gets registered, and it matters which one you pick:

Benzene.HealthChecks.Core adds these builder extensions on top of the two core methods:

builder.AddHealthCheck(IHealthCheck healthCheck);                      // wraps an existing instance
builder.AddHealthChecks(params IHealthCheck[] healthChecks);           // several existing instances at once
builder.AddHealthCheckFactory(IHealthCheckFactory healthCheckFactory); // e.g. HttpPingHealthCheckFactory, DatabaseHealthCheckFactory<T>

Benzene.HealthChecks adds inline overloads for one-off checks that don't need their own class, covering sync/async and bool/IHealthCheckResult return types, with or without an explicit Type:

builder.AddHealthCheck("my-check", resolver => true);
builder.AddHealthCheck("my-check", async resolver => await SomeAsyncCheckAsync());
builder.AddHealthCheck("my-check", resolver => HealthCheckResult.CreateWarning("my-check"));
builder.AddHealthCheck(resolver => true);                              // Type defaults to "inline"

All of the inline overloads build an InlineHealthCheck under the hood — you rarely construct InlineHealthCheck yourself.

Built-in health checks

SimpleHealthCheck (Benzene.HealthChecks)

Always succeeds. Type is "Simple". Useful as a smoke test or a placeholder while you build out real checks.

InlineHealthCheck (Benzene.HealthChecks)

Wraps a Func<Task<IHealthCheckResult>> with an optional explicit Type (defaults to string.Empty, which the response namer then names "HealthCheck-1" — the bare "HealthCheck" key is pre-reserved, so an empty type never lands on it; see naming). This is what every inline AddHealthCheck(...) overload above constructs for you.

FailedHealthCheck (Benzene.HealthChecks)

Wraps an Exception. Type is "Failed"; ExecuteAsync() always returns a Failed result with the exception's message in Data["Exception"]. Extensions.BuildHealthCheck(Func<IHealthCheck>) uses this to turn a check construction failure (e.g. a factory that throws) into a reportable result instead of an unhandled exception.

MemoryHealthCheck (Benzene.HealthChecks)

A host self-check on this process's memory. Type is "Memory"; it measures the process working set (Environment.WorkingSet — the physical memory a container/host OOM-killer watches) and reports Failed at or above a hard ceiling, an optional Warning at or above a soft ceiling (degraded but not fatal — does not flip aggregate IsHealthy), otherwise healthy. Data includes WorkingSetBytes, ManagedHeapBytes (GC.GetTotalMemory), MaximumBytes, and the GC-reported GcTotalAvailableBytes/GcHighLoadThresholdBytes for diagnostics.

using Benzene.HealthChecks;

app.UseHealthCheck("healthcheck", x => x
    .AddMemoryCheck(maximumBytes: 900 * 1024 * 1024, warningBytes: 700 * 1024 * 1024));

Set the ceiling below the container/pod memory limit so a readiness probe flips before the runtime is OOM-killed. It has no external dependency (only Environment/GC), so unlike DiskHealthCheck it ships in Benzene.HealthChecks directly rather than a dedicated package.

ShutdownReadinessHealthCheck (Benzene.HealthChecks) — graceful drain

Fails a readiness probe once graceful shutdown begins, so a load balancer / Kubernetes stops routing new traffic to the instance while its in-flight requests finish, before the process exits. Type is "Shutdown"; Data["ShuttingDown"] is the latch state.

It reads a caller-owned ShutdownState latch. Wire the latch to your host's shutdown signal with LinkTo(CancellationToken) — on the generic host / ASP.NET Core that's IHostApplicationLifetime.ApplicationStopping (which fires on SIGTERM):

using Benzene.HealthChecks;

var shutdown = new ShutdownState()
    .LinkTo(app.Services.GetRequiredService<IHostApplicationLifetime>().ApplicationStopping);

app.UseBenzene(benzene => benzene
    .UseHttp(http => http
        .UseReadinessCheck(x => x.AddShutdownReadinessCheck(shutdown))
        .UseLivenessCheck() // deliberately NOT the shutdown check — a draining instance is alive
        .UseMessageHandlers()));

Add it only to readiness, never liveness: a draining instance is healthy (don't restart it), just not ready for new traffic. On a host without an ApplicationStopping token you can trip the latch yourself (shutdown.MarkShuttingDown()) from a PosixSignalRegistration for SIGTERM or any shutdown hook. Dependency-free (only a BCL CancellationToken), so it also ships in Benzene.HealthChecks directly.

TimeOutHealthCheck / ExceptionHandlingHealthCheck (internal safety net)

These two are internal — you never construct them yourself, but it's worth knowing they run around every check automatically. HealthCheckProcessor.PerformHealthChecksAsync wraps each IHealthCheck in ExceptionHandlingHealthCheck (catches any exception thrown by ExecuteAsync() and turns it into a Failed result with Data["Exception"]) and then TimeOutHealthCheck (a hard-coded 10 second timeout — if the check hasn't completed by then, it returns a Failed result with Data["Error"] = "Timed Out" instead of waiting indefinitely). Neither timeout nor exception handling is currently configurable; every check gets the same 10 second budget and the same catch-all.

HttpPingHealthCheck (Benzene.HealthChecks.Http)

GETs a URL and reports healthy only on a 200 OK response (any other status code, including other 2xx codes, is Failed). Type is "HttpPing"; Data includes Url and StatusCode.

using Benzene.HealthChecks.Http;

app.UseHealthCheck("healthcheck", x => x
    .AddHttpPing("https://downstream-service/health"));

AddHttpPing(url) registers an HttpPingHealthCheckFactory, which resolves a plain HttpClient from the resolver at check time — make sure one is registered (e.g. via services.AddHttpClient()/IHttpClientFactory, or a manual HttpClient registration).

Entity Framework Core health checks (Benzene.HealthChecks.EntityFramework)

Two checks, both generic over your DbContext type and resolving it via constructor injection — neither has a dedicated .AddHealthCheck<T>()-style extension method, so register them either as a DI-resolved type or via their factory:

Wiring into the pipeline: .UseHealthCheck()

Message-topic based (every transport)

Defined in Benzene.HealthChecks.Extensions, generic over TContext, so it works on any pipeline (BenzeneMessage, SNS, SQS, Kafka, Azure Functions, ASP.NET Core, self-hosted workers):

app.UseHealthCheck(topic, params IHealthCheck[] healthChecks);
app.UseHealthCheck(topic, Action<IHealthCheckBuilder> action);
app.UseHealthCheck(topic, IHealthCheckBuilder builder);

The middleware checks the incoming message's topic against both the topic you passed in and the constant default topic "healthcheck" (Constants.DefaultHealthCheckTopic) — so even if you wire it up under a custom topic like "orders-service:healthcheck", it's still reachable via the plain "healthcheck" topic too. If neither matches, the middleware calls next() and gets out of the way.

app.UseHealthCheck("orders-service:healthcheck", x => x
    .AddHealthCheck<SimpleHealthCheck>()
    .AddHealthCheckFactory(new DatabaseHealthCheckFactory<OrdersDbContext>("20220809094008_V8")));

HTTP method + path based (raw HTTP transports)

Two packages that deal with raw HTTP requests directly — Benzene.SelfHost.Http and Benzene.Aws.Lambda.ApiGateway — additionally expose UseHealthCheck overloads that match on HTTP verb and path instead of (or in addition to) a message topic, since those pipelines see the raw request before any topic has been resolved from it:

// Benzene.SelfHost.Http / Benzene.Aws.Lambda.ApiGateway
app.UseHealthCheck(method, path, params IHealthCheck[] healthChecks);
app.UseHealthCheck(topic, method, path, params IHealthCheck[] healthChecks);
app.UseHealthCheck(method, path, Action<IHealthCheckBuilder> action);
app.UseHealthCheck(topic, method, path, Action<IHealthCheckBuilder> action);
app.UseHealthCheck(topic, method, path, IHealthCheckBuilder builder);
apiGatewayApp.UseHealthCheck("healthcheck", "POST", "/healthcheck", healthChecks);

These match when the request's HTTP method and path equal method/path exactly (method comparison is case-insensitive); when they don't match, the middleware falls through to next() the same way the topic-based version does. ASP.NET Core doesn't currently have an equivalent method+path overload — use the topic-based .UseHealthCheck(topic, ...) there (an ASP.NET Core request's topic is already resolved via routing before it reaches this middleware, so there's no raw path left to match on at this point in the pipeline).

Kubernetes-style liveness/readiness: .UseLivenessCheck() / .UseReadinessCheck()

.UseHealthCheck() runs one undifferentiated set of checks. For Kubernetes (or any platform with a liveness-vs-readiness distinction), Benzene.HealthChecks also provides two purpose-built convenience methods — see Kubernetes Health Checks for the full guide, including which checks belong in which and example probe YAML. Short version:

// Benzene.HealthChecks — topic-based, every transport
app.UseLivenessCheck(x => x.AddHealthCheck<ProcessResponsiveCheck>());
app.UseReadinessCheck(x => x.AddHealthCheckFactory(new DatabaseHealthCheckFactory<MyDbContext>("...")));

Unlike .UseHealthCheck(), these do not also respond to Constants.DefaultHealthCheckTopic — each only matches its own topic (Constants.DefaultLivenessTopic/DefaultReadinessTopic), so registering both in the same pipeline doesn't have one silently shadow the other on a shared fallback topic.

When a Benzene client is configured (e.g. .UseSqs(queueUrl)), it auto-registers a non-destructive reachability check for that dependency on the general healthcheck topic only — never a liveness or readiness probe, because a shared-downstream check gating a probe is a cascading-failure risk. Opt out per client with healthCheck: false. See Kubernetes Health Checks for the full reasoning.

There is a third purpose-built method, .UseContractsCheck(), that answers a probe-less contracts diagnostic topic (Constants.DefaultContractsTopic) — for consumer-side contract-drift checks (Benzene.Clients.HealthChecks' ClientHealthCheck / AddContractCheck<TClient>(serviceName)) that call a downstream service. These belong in neither liveness nor readiness — a struggling downstream would otherwise restart or de-route otherwise-healthy pods — so monitoring/the mesh scrape contracts directly. See Kubernetes Health Checks.

Benzene.SelfHost.Http and Benzene.Aws.Lambda.ApiGateway additionally expose HTTP-path versions, defaulting to the conventional Kubernetes probe paths:

// Benzene.SelfHost.Http / Benzene.Aws.Lambda.ApiGateway
app.UseLivenessCheck(checks);   // GET /livez
app.UseReadinessCheck(checks);  // GET /readyz
app.UseLivenessCheck("/custom/live/path", checks);  // path override

Both variants (topic-based and HTTP-path-based) run through the same HealthCheckProcessor, so an unhealthy result maps to HTTP 503 the same way .UseHealthCheck()'s does — see HTTP status codes below.

gRPC (grpc.health.v1)

Package: Benzene.Grpc.AspNet. This one is different in kind from everything above: instead of routing a message topic or HTTP path through .UseHealthCheck(), it bridges Benzene's health checks onto the standard grpc.health.v1 protocol, so any generic gRPC health-checking tool (grpc_health_probe, Kubernetes gRPC liveness probes, grpcurl) can query it without knowing anything about Benzene:

services.AddBenzeneGrpc(o => o.EnableHealthChecks = true);
services.AddScoped<IHealthCheck, DatabaseHealthCheck>();   // Benzene.HealthChecks.Core.IHealthCheck

app.MapGrpcService<GreeterService>();
app.MapBenzeneGrpcHealthService();

BenzeneHealthCheckBridge is an ASP.NET Core Microsoft.Extensions.Diagnostics.HealthChecks.IHealthCheck (not the same interface as above — note the different namespace) registered as the "benzene" check; it resolves every Benzene.HealthChecks.Core.IHealthCheck from the container directly (plain services.AddScoped<IHealthCheck, T>(), no IHealthCheckBuilder involved) and aggregates them: unhealthy if any failed, degraded if any warned, healthy otherwise. A gRPC Check/Watch call then reflects that aggregate as SERVING/NOT_SERVING per the standard protocol.

To split liveness from readiness over gRPC (so a liveness probe doesn't run your external-dependency checks), set LivenessCheckTypes/ReadinessCheckTypes on BenzeneGrpcOptions — Benzene then publishes named "liveness"/"readiness" grpc.health.v1 services reporting only those check Types, alongside the overall "" service. See Kubernetes Health Checks.

Both health checks and reflection are off by default (EnableHealthChecks/EnableReflection on BenzeneGrpcOptions) — see gRPC Setup for the full walkthrough.

Response format

The response payload is a HealthCheckResponse (IHealthCheckResponse<HealthCheckResult>), serialized through whatever ISerializer your app uses. By default that's a camelCase naming policy, which applies to the actual C# properties on HealthCheckResponse/HealthCheckResult (IsHealthy, HealthChecks, Status, Type, Data) but not to the contents of the Data dictionary itself, since a naming policy only rewrites declared properties, not arbitrary IDictionary<string, object> keys. For example:

{
  "isHealthy": true,
  "healthChecks": {
    "Simple": { "status": "ok", "type": "Simple", "data": {} },
    "Database": {
      "status": "ok",
      "type": "Database",
      "data": {
        "CanConnect": true,
        "AppliedMigrations": ["20220101000000_Init", "20220809094008_V8"],
        "TargetMigration": "20220809094008_V8",
        "MigrationMatch": true,
        "MigrationContains": true,
        "Error": null
      }
    }
  }
}

That's why data's inner keys stay PascalCase (CanConnect, AppliedMigrations, ...) exactly as the built-in checks set them, while everything around them follows your serializer's naming policy. This example assumes the default System.Text.Json-based ISerializer (Benzene.Core.MessageHandlers), whose PropertyNamingPolicy doesn't touch dictionary contents; if your app instead wires up Benzene.NewtonsoftJson's serializer, check its CamelCasePropertyNamesContractResolver configuration — Newtonsoft's camel-casing can be configured to rewrite dictionary keys too, which would camelCase Data's contents as well.

If a check's status is ok/warning, isHealthy stays true; a single failed check flips the whole response to false even if every other check succeeded.

HTTP status codes

Over an HTTP transport, the response's HTTP status code reflects isHealthy, not just its body: 200 when healthy, 503 Service Unavailable when not (mapped by DefaultHttpStatusCodeMapper, same as everywhere else in Benzene). This matters because most health-check consumers — a Kubernetes HTTP probe, a load balancer target-health check — only look at the status code, not the response body, so a check whose body says "isHealthy": false but whose status code is still 200 would be silently treated as healthy by those consumers. HealthCheckProcessor.PerformHealthChecksAsync handles this for every .UseHealthCheck()/.UseLivenessCheck()/.UseReadinessCheck() variant, over every transport, with no extra configuration needed.

Data["Error"]/Data["Exception"] entries (populated when a check fails with an exception, or times out) report the exception's type name (e.g. "SqlException"), not its message — some ADO.NET providers embed connection details in exception messages, and this response can flow out to whatever calls the health check with no built-in authorization. See Privacy & Data Handling for the full reasoning.

Result naming and deduplication

Each entry's key in healthChecks comes from HealthCheckNamer, run against the Type on the executed HealthCheckResult — not necessarily the IHealthCheck instance's own Type property:

Example: a custom health check class

The cleanest way to add a non-trivial check is as its own class, which is also the shape AddHealthCheck<T>()/DI resolution expects:

using Benzene.HealthChecks.Core;
using Microsoft.EntityFrameworkCore;

public class DatabaseConnectionHealthCheck<TDbContext> : IHealthCheck where TDbContext : DbContext
{
    private readonly TDbContext _dbContext;
    public string Type => "DatabaseConnection";

    public DatabaseConnectionHealthCheck(TDbContext dbContext)
    {
        _dbContext = dbContext;
    }

    public async Task<IHealthCheckResult> ExecuteAsync()
    {
        var canConnect = await TryConnect(_dbContext);
        return HealthCheckResult.CreateInstance(canConnect, Type, new Dictionary<string, object>
        {
            { "CanConnect", canConnect },
        });
    }

    private static async Task<bool> TryConnect(DbContext dbContext)
    {
        try
        {
            return await dbContext.Database.CanConnectAsync();
        }
        catch
        {
            return false;
        }
    }
}

(This is, in fact, exactly what Benzene.HealthChecks.EntityFramework's own DatabaseConnectionHealthCheck<TDbContext> looks like — you only need to write your own version of this if you want different Data/behavior than the shipped one.)

Troubleshooting

See Also