Observability in 2026: OpenTelemetry Graduated, the Collector Is Still v0.162, eBPF Instrumentation Is v0.14, and Datadog Bills $1.12B a Quarter

Oct 5, 2026 · 13 min · Ajay Kumar

Every microVM on my platform emits telemetry twice: once from the host, where I own the kernel and can see every syscall a tenant makes, and once from inside the guest, where the customer's own SDK decides what to record. (PandaStack is my company, so read what follows knowing that.) The host side is cheap and complete and tells me nothing about the application; the guest side is expensive and partial and tells me everything. The observability argument of 2026 is a version of that split: how much evidence to collect from outside the process with eBPF, how much to instrument by hand, and how to stop the result becoming the second-largest line on the cloud bill. From the release notes, SEC filings, pricing pages and surveys, this is what is actually stable, what is actually cheap, and where the two overlap.

Graduated is not the same as finished

OpenTelemetry became a CNCF graduated project on 21 May 2026. The announcement counts over 12,000 contributors from over 2,800 companies, the second-highest velocity of any CNCF project after Kubernetes, and 1.36 billion downloads of the JavaScript API package, with Python past 1.3 billion. Traces, metrics and logs are stable at the API, SDK and wire level; OTLP is at 1.11.0, stable for those three signals and Development for profiles.

The Collector is where the word "stable" needs qualifying. The core repository tags each release twice, and the latest is v1.68.0/v0.162.0 on 28 September 2026: the 1.x number belongs to modules that have individually passed a stabilisation issue, and the 0.x number is everything else, on a two-week release cycle with no distribution-wide 1.0 criteria written down. The documentation calls the Collector's overall status "mixed" and tells you to read each component's README. The contrib repository started its own stable line this September, v1.0.0/v0.161.0 on 15 September and then v1.1.0/v0.162.0, carried by the Kubernetes attributes processor reaching v1.0.0 on 16 September. The otelcol-contrib 0.162.0 manifest still pins the OTLP receiver, the batch processor and the OTLP exporter at v0.162.0. If someone tells you "the collector went GA in 2026", they mean the project graduated. The binary you run did not.

Signal or component Status, October 2026 Evidence
Traces, metrics, logs API/SDK and OTLP Stable OTLP spec 1.11.0
Profiles signal (OTLP, eBPF profiler) Alpha, "not for critical production workloads" Profiles alpha post, 26 Mar 2026
Collector core distribution v1.68.0/v0.162.0, overall status "mixed" Releases, 28 Sep 2026
k8sattributes processor v1.0.0 (first contrib stable module line) Blog, 16 Sep 2026
tail_sampling processor Beta, traces only README
transform processor Beta (traces, metrics, logs); Development for profiles README
filter processor Alpha (traces, metrics, logs) README
OTel Arrow exporter Beta, traces, metrics, logs README
Semconv: HTTP Stable since v1.23.0 Blog, Nov 2023
Semconv: database Stable since May 2025 RPC stabilisation post
Semconv: Kubernetes resource attributes Stable in v1.42.0; K8s metrics release candidate v1.42.0, 12 Jun 2026
Semconv: messaging, RPC Development (RPC stabilisation project since June 2025) Spec page
OBI (eBPF instrumentation) v0.14.0, "currently in Development" Releases, 2 Oct 2026

The fourth signal arrived as an alpha

Profiling became an official signal in public alpha on 26 March 2026: an OTLP profiles format that the post says is about 40 percent smaller on the wire than the previous draft; Collector support from v0.148.0 including a pprof receiver and OTTL for profiles; and the eBPF profiler as an official collector component. That profiler is the agent Elastic pledged in early 2024. What remains before stable is listed honestly: a symbolisation API, process and thread context sharing so a profile can be joined to a span, and backends; the SIG says not to use it for critical production workloads. The profiler needs no SDK at all, so this is the first signal where collecting from outside the process is the default rather than the fallback.

Conventions: HTTP in 2023, databases in 2025, queues still moving

Semantic conventions decide whether your dashboards survive an SDK upgrade. The registry is at v1.44.0 as of 4 August 2026. HTTP was declared stable first, in v1.23.0 in November 2023; database followed in May 2025, which is why db.system.name, db.namespace, db.operation.name and db.query.text are now Stable and db.connection_string and db.user were removed rather than renamed. Kubernetes resource attributes went to release candidate in March and stable in June's v1.42.0, the same release that moved every gen_ai.* attribute to its own repository; I covered those in the agent observability post and will not repeat it.

Messaging is still Development, and RPC's stabilisation project only started in June 2025 with no date attached. So a Kafka or gRPC span from one SDK may not match one from another, and OTEL_SEMCONV_STABILITY_OPT_IN, which lets a library emit old, new or both attribute sets, will be in your manifests for another year or two. The discipline is boring and necessary: pin SDK versions, run dual-emit across a release, diff the attribute keys before you flip dashboards. No assistant will notice that http.status_code became http.response.status_code until the alert that depended on it stops firing.

eBPF: one upstream, several distributions, no 1.0

Grafana Labs donated Beyla to OpenTelemetry in May 2025, and the result, OpenTelemetry eBPF Instrumentation, shipped its first alpha on 3 November 2025 with Grafana Labs, Splunk, Coralogix and Odigos contributing. It instruments at the protocol layer, so it covers HTTP/1.1 and 2, gRPC, SQL, Redis, MongoDB, Kafka and S3 for any language, and it is candid about limits: context propagation works for Go, Node.js, Python, NGINX and PHP, and not for reactive frameworks, Java virtual threads or complex thread pools. The project's 2026 goals, published 23 January, are a stable 1.0, MQTT, AMQP and NATS, .NET including the 4.x Framework, and a hybrid mode where OBI and SDK telemetry carry consistent labels.

Eleven releases later it is at v0.14.0, tagged 2 October 2026, and the notes show what zero-code instrumentation has become: probes attach without CAP_SYS_ADMIN where tracefs allows, GPU metrics, route templates harvested from Django, FastAPI, Flask, Rails, .NET, Symfony and Laravel so that /users/{id} is one series rather than a million, and per-stream traceparents for multiplexed HTTP/2 and gRPC. The README still says "OBI is currently in Development": expect breaking changes between minor versions, pin a tag, do not assume telemetry continuity for dashboards. Grafana's Beyla 2.5 is now explicitly a downstream distribution of OBI.

The alternatives have settled into roles. Odigos, Apache-2.0 at v1.38.0 this week, uses eBPF only for Go and injects the native OpenTelemetry agents for Java, Python, .NET and Node.js. Coroot, Apache-2.0 at v1.27.0, collects all four signals with eBPF into ClickHouse and Prometheus and charges $1 per monitored CPU core per month for Standard. Pixie, New Relic's donation and a CNCF sandbox project since June 2021, remains a Kubernetes debugger with full-body capture at a claimed under 5 percent of cluster CPU, exporting through PxL scripts to a collector. A fourth path is not eBPF at all: OpenTelemetry's Go compile-time instrumentation, built with Alibaba and Datadog, reached v1.0.0 on 28 September and injects instrumentation through -toolexec at build time. The summary: eBPF gives you RED metrics and service maps for everything, including binaries you cannot rebuild; SDKs give you the business attributes and the context propagation eBPF cannot see. You will run both, and consistent labels between them is the 2026 goal that matters.

What the bill looks like

Datadog's second quarter, in its 6 August 8-K, was $1.12 billion of revenue, up 36 percent, with about 4,720 customers paying $100,000 or more a year, up 23 percent, and full-year guidance of $4.45 to 4.47 billion. Nothing in the growth suggests the market is tired of paying; everything in the surveys suggests the people doing the paying are.

The reference story is still Coinbase. On Datadog's 4 May 2023 earnings call the CFO referred to a large upfront bill from a crypto customer in the first quarter of 2022; analysts put it at about $65 million, and the Pragmatic Engineer confirmed through Coinbase engineers that it covered 2021 usage. The useful part is the sequel: Coinbase stood up a Grafana, Prometheus and ClickHouse stack, double-wrote for months, then stayed on Datadog after Datadog cut the price. The open-source stack is a negotiating instrument before it is a migration.

List prices explain how a bill gets there. From Datadog's pricing page: infrastructure $15 per host per month on Pro and $23 on Enterprise; APM from $31 per host; logs at $0.10 per GB ingested plus $1.70 per million events indexed for 3 to 30 days, or $0.05 per million events parked in Flex storage. Ingest is cheap and indexing is where the money is, which is why every cost guide begins with "index less".

Vendor, October 2026 Unit price (list, annual where offered) Licence or model Source
Datadog Log Management $0.10/GB ingested + $1.70 per M events indexed (15 d); Flex $0.05 per M stored SaaS Pricing
Grafana Cloud logs and traces $0.05/GB process + $0.40/GB write + $0.10/GB-month retain; metrics $6.50 per 1k series; free tier 50 GB logs, 50 GB traces, 10k series SaaS; Loki is AGPL-3.0 Pricing
Honeycomb Free to 20M events/month; Pro from $150/month to 750M events; Enterprise from 10B events/year SaaS, per event Pricing
Dash0 $0.06 per M spans or logs ingested + $0.54 per M stored 30 d; metrics $0.02 + $0.18 per M points SaaS, OTel-native Pricing
SigNoz Cloud $0.30/GB logs or traces, 15 d; $0.10 per M metric samples; $49/month base SaaS; self-host OSS, v0.144.0 Pricing
OpenObserve Cloud $0.50/GB ingested (after 30% annual discount) + $0.01/GB queried SaaS; self-host AGPL-3.0, v1.0.4 Pricing
Chronosphere No public price; pricing page returns 404 Sales-led Checked 5 Oct 2026

The surveys agree on the shape. Grafana Labs' 2026 Observability Survey, 1,363 respondents in 76 countries between October 2025 and January 2026, ranks complexity and overhead as the top concern at 38 percent, signal-to-noise at 34 and cost at 31; half now use SaaS for observability, up from 43 percent, and half expect to spend more next year. The 2025 edition had 37 percent calling costs too high and about eight tools per organisation. The counterweight is adoption: 76 percent have invested in OpenTelemetry, 41 percent run it in production, 57 percent for metrics, 50 for traces and 48 for logs, citing ease of adoption (41 percent) and freedom to switch vendors (37).

Wide events, columnar storage and who is selling what

The intellectual case against the bill is Charity Majors' "observability 2.0" from August 2024: store each request once, as one arbitrarily wide structured event aggregated at query time, instead of several times as metric, log line, span and APM record that cannot be correlated. Her cost-crisis piece puts numbers on the failure mode: paying to store telemetry five ways, bills growing three to ten times faster than traffic, a single custom metric costing $30,000 a month in one case. It is a Honeycomb pitch, and it is also correct, and since 2024 the storage engines that make wide events cheap have become commodity.

ClickHouse bought HyperDX on 13 March 2025 and packaged it as ClickStack: ClickHouse and the OTel collector under Apache-2.0, the HyperDX UI under MIT. ClickHouse's use-case page claims 10x to 30x compression on OTel data and "up to 90 percent" less storage, lists Anthropic, Netflix and Cloudflare as users, and cites a 200x cost reduction against Datadog from its own migration; vendor numbers, reported as such. OpenObserve, Rust and Parquet on object storage, hit 1.0 under AGPL-3.0. VictoriaLogs, Apache-2.0 at v1.53.0 on 1 October, claims up to 30 times less RAM and 15 times less disk than Elasticsearch and Loki, again a vendor benchmark. Grafana Loki 3.0 in April 2024 added native OTLP ingestion, structured metadata and experimental bloom filters; it is at v3.7.8 under AGPL-3.0, and since 16 March 2026 the community Helm chart lives in a grafana-community fork. Quesma, once a gateway for putting ClickHouse behind Kibana, now markets cost analytics for AI coding agents, which tells you where the easier money went.

From instrumentation to invoice: what is stable, and where the bytes get cheaper 1. Sources SDK traces, metrics, logs: stable OBI eBPF v0.14.0: Development eBPF profiler: profiles alpha Go compile-time instr.: v1.0.0 Lever: metric cardinality limit, default 2,000 streams per metric 2. Collector v0.162.0 (mixed) filter (alpha): drop health checks transform (beta): delete, truncate tail_sampling (beta): errors, slow OTel Arrow exporter (beta) Lever: ~50% less bandwidth than OTLP/gRPC + zstd (Arrow README) 3. Storage and bill Columnar OSS: ClickHouse, Loki, VictoriaLogs, SigNoz, OpenObserve Datadog logs: $0.10/GB in + $1.70 per M events indexed, 15 d Grafana Cloud logs: $0.45/GB + $0.10 per GB-month retained Lever: index less, retain less, aggregate at query time Grafana Labs survey 2026 (n=1,363): top concerns complexity 38%, signal-to-noise 34%, cost 31% Datadog Q2 2026 revenue $1.12B, up 36%; the Coinbase bill for 2021 was $65M (Datadog call, May 2023)
Stability labels from component READMEs and the OTel blog as of 5 October 2026; list prices from vendor pricing pages the same day; survey figures from Grafana Labs' 2026 report; Datadog revenue from its Q2 2026 8-K.

The collector is where the invoice is decided

Whatever the backend, the bytes that reach it were shaped by a pipeline you own, which has four cost controls mature enough to rely on. The SDK-side metric cardinality limit, which the spec sets at a default of 2,000 attribute combinations per metric stream, folds overflow into one series tagged otel.metric.overflow=true, a signal to inspect rather than a limit to raise. OTTL in the filter and transform processors drops and rewrites telemetry before it is billed. Tail sampling keeps every error and every slow trace and a few percent of the rest, with one hard constraint: every span of a trace must reach the same collector instance, so you run a stateless front tier using the load_balancing exporter with routing_key: traceID and a sampling tier behind it. And transport, if you pay for egress: the OTel Arrow exporter's README expects about 50 percent less bandwidth than OTLP/gRPC with zstd at equal batch sizes, and the June Phase 2 post reports the Rust OTAP dataflow engine at 10 to 20 times the throughput of the Go OTLP path, while labelling it incubation-stage with no stability guarantee.

# otelcol-contrib 0.162.0 (contrib v1.1.0/v0.162.0, 29 Sep 2026). Sampling tier.
# A front tier of collectors runs the load_balancing exporter with routing_key: traceID
# pointing at this tier's headless Service, so every span of a trace lands on one instance.
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  memory_limiter:                 # always first; refuses data instead of OOM-killing the pod
    check_interval: 1s
    limit_mib: 1800
    spike_limit_mib: 400

  filter/noise:                   # alpha for traces/metrics/logs: drop what nobody will query
    error_mode: ignore
    trace_conditions:
      - span.attributes["url.path"] == "/healthz" or span.attributes["url.path"] == "/readyz"
      - span.attributes["http.request.method"] == "OPTIONS"
    metric_conditions:
      - resource.attributes["deployment.environment.name"] == "dev" and metric.name == "http.server.active_requests"

  transform/trim:                 # beta: fix cardinality and secrets before they are indexed
    error_mode: ignore
    trace_statements:
      - delete_key(span.attributes, "http.request.header.authorization")
      - replace_pattern(span.attributes["url.full"], "token=[^&]*", "token=REDACTED")
      - replace_pattern(span.attributes["url.path"], "/[0-9a-f]{8}-[0-9a-f-]{27}", "/{uuid}")
      - truncate_all(span.attributes, 1024)
    log_statements:
      - keep_keys(resource.attributes, ["service.name", "service.namespace", "k8s.namespace.name", "k8s.pod.name", "deployment.environment.name"])

  tail_sampling:                  # beta, traces only. Decisions are per trace after decision_wait.
    decision_wait: 10s
    num_traces: 200000            # traces held in memory; size for decision_wait x arrival rate
    expected_new_traces_per_sec: 2000
    policies:
      - name: errors              # keep every trace with an ERROR span
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow                # keep every trace over 2 s
        type: latency
        latency: {threshold_ms: 2000}
      - name: checkout-5xx        # keep server errors on a service you care about, by OTTL
        type: ottl_condition
        ottl_condition:
          error_mode: ignore
          span:
            - 'resource.attributes["service.name"] == "checkout" and span.attributes["http.response.status_code"] >= 500'
      - name: baseline            # and 5 percent of everything else, for the dashboards
        type: probabilistic
        probabilistic: {sampling_percentage: 5}

  batch:
    send_batch_size: 8192
    timeout: 5s

exporters:
  otlp/backend:                   # any OTLP backend: ClickStack, Grafana, Dash0, Datadog, ...
    endpoint: otel-gateway.observability.svc:4317
    compression: zstd
    retry_on_failure:
      enabled: true
    sending_queue:
      enabled: true
      num_consumers: 8
      queue_size: 5000

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, filter/noise, transform/trim, tail_sampling, batch]
      exporters: [otlp/backend]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, filter/noise, batch]
      exporters: [otlp/backend]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, transform/trim, batch]
      exporters: [otlp/backend]
# Front tier (same 0.162.0 image): no sampling here, just consistent routing by trace ID.
exporters:
  load_balancing:
    routing_key: traceID
    protocol:
      otlp:
        compression: zstd
        tls: {insecure: true}        # in-cluster; terminate TLS at the sampling tier if you prefer
    resolver:
      dns:
        hostname: otel-sampler-headless.observability.svc
        port: 4317

Two cautions from the READMEs: the filter processor is alpha, so pin the image and read the changelog every two weeks; and num_traces is a memory budget, not a target, so at 2,000 new traces a second and a 10-second wait you need at least 20,000 in flight plus burst headroom.

What I take from all of it. The graduation is real and the data model is settled where it matters: traces, metrics, logs, HTTP, databases and Kubernetes attributes are stable. Everything that would let you stop writing instrumentation, OBI, the profiler, the Arrow engine, is at v0 and says so on the tin; I would run OBI today for RED metrics and service maps and keep SDKs for the attributes that pay the bills. The cost problem is not a vendor problem. Datadog grows 36 percent a year because ingesting everything and indexing it is the default, and the alternative is not a cheaper vendor but a pipeline with a cardinality limit, a filter, a redaction pass and a tail sampler in front of columnar storage, which is work a platform team has to own. This quarter I am pinning my collectors to one contrib release, folding the per-request log lines my guests emit into one wide event per request before storage, and putting the eBPF profiler on a few hosts as the alpha its authors call it. For what tenants do inside their VMs, which I wrote about in the 2 AM post, the host-side record stays the source of truth: the one signal nobody can turn off from inside the guest, and the one I never have to sample.


Related: Agent Observability in 2026: The Spans Are Standard, the Standard Isn't Stable, It's 2 AM. Do You Know What Your AI Agent Is Doing? and FinOps for AI Token Costs in 2026.

I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.

Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.


Related

Oct 5, 2026

AIOps in 2026: 47% on ITBench, 10% on Hard On-Call, $6.50 an Investigation, and Four Outages Where the Automation Was the Incident

The incident-response agents of 2026 against the evidence: Datadog Bits AI SRE GA at about 6.5 credits ($6.50) per investigation, PagerDuty's approval-gated SRE Agent, Grafana's six GA agent tools, Splunk's AI SRE, Elastic buying Deductive, Resolve at a $1B headline valuation on $4M ARR, Traversal's $48M; benchmarks from 13.8% on ITBench in 2025 to 47% in May 2026, 20.7% exact root cause on OpenRCA 2.0, 10% on hard ORCA-Bench and 40% hallucinated causes; what the AWS, Azure, Cloudflare and Google postmortems say about automation; and why a human should hold the button.

14 min
Oct 5, 2026

DevOps Still Matters in 2026: AI Cut Delivery Stability 7.2%, Then Doubled Merged PRs and Added 91% to Review Time

Why DevOps is the big thing of the agent era: DORA 2024 found a 25% rise in AI adoption cost 1.5% throughput and 7.2% stability, DORA 2025 saw throughput turn positive while instability stayed up across nearly 5,000 respondents, and DORA's 2026 ROI model budgets a 15% three-month dip and a change failure rate rising from 5% to 6%; GitHub merged 518.7M PRs (+29%) and over 1M agent PRs in five months, Faros telemetry on 10,000 developers shows 98% more PRs and 91% longer reviews, METR found experienced developers 19% slower, and the Replit postmortem's fixes are 2015 DevOps controls.

13 min
Oct 5, 2026

FinOps for AI in 2026: 98% Manage Token Spend, 3 in 4 Can't Prove the Value, GPUs Run at 5%, and FOCUS Gets a Token Column in 1.5

The AI cost ledger as of October 2026: State of FinOps 2026 (1,192 practitioners, 98 percent manage AI spend, up from 31 percent in 2024), the Tokenomics Foundation finding that three in four enterprises cannot prove AI outcomes to the CFO, a per-million-token price table for Anthropic, OpenAI, Google and DeepSeek with cache and batch multipliers, a 20-step agent loop at $1.58 uncached and $0.42 cached, the FOCUS 1.2 to 1.5 roadmap and who exports which version, LiteLLM and Cloudflare budgets, Cast AI's 5 percent GPU utilisation, and $0.0037 of model cost inside a $0.99 resolution.

13 min