What changed in early September
On 2 September 2026, TechCrunch reported that HiddenLayer had raised $100 million to help enterprises secure AI deployments. The report framed AI security less as a model-only problem and more as an infrastructure control-plane problem: discovery, identity, policy controls and governance around deployed AI systems.
In the week of 7-13 September 2026, Product Hunt’s infrastructure leaderboard also surfaced two adjacent tools: AI Observability by OpenObserve, described as OpenTelemetry-native observability for agents and LLMs, and Harden, a security layer for AI coding agents. These are not the same product category, but they point to the same operational shift. AI agents are becoming normal workloads, and they bring traces, prompts, tool calls, generated code, data access, credentials and cost allocation into the cluster operator’s scope.
For a hyperscaler, the answer is another managed service, another identity layer and another bill. For an operator with a quarter rack, one rack or a few racks of owned hardware, the answer has to be more direct. The same GPU estate must run bare-metal training jobs, Kubernetes inference, VM tenants, notebooks, coding agents and shared storage. The operator must know who used what, which GPU slice ran which job, which tenant generated which trace, and whether policy allowed that workload to reach the data, the internet or another tenant’s service.
That is the real bottleneck behind the funding and the Product Hunt launches: once multiple teams share a small-to-mid GPU fleet, AI observability and AI security become multi-tenant infrastructure problems. They cannot be bolted on only at the application layer.
The bottleneck operators will hit first
A small GPU estate usually starts with a simple pattern. One team needs model fine-tuning, another wants inference, a third wants agentic coding tools, and a fourth wants a VM for a vendor appliance. The first month works because people coordinate in chat. The second month brings missed reservations, idle GPUs, untagged workloads and unclear ownership of storage. By the third month, finance asks for chargeback, security asks which service account accessed restricted data, and facilities asks why power draw is high while the business sees a GPU shortage.
AI agents make this worse because their unit of work is not only a container or a notebook. An agent can call tools, open pull requests, query a vector database, invoke an LLM endpoint, write logs containing sensitive context, and run multiple short tasks over a long session. Observability has to cover the chain. Security has to know the identity behind it. Metering has to attribute GPU time, CPU time, storage, egress and possibly model tokens to the right tenant.
The failure mode is familiar to cluster operators: shared infrastructure with application-level ownership but no enforceable infrastructure boundary. Teams may deploy their own tracing, secrets and agent security tools, but the GPU fleet still lacks a consistent tenancy model. A pod can see too much. A VM consumes a whole accelerator while using 20 percent of it. A fine-tuning job leaves 12 TB of checkpoints in shared storage. An inference endpoint consumes capacity meant for a paying tenant. Logs leave the country because the default SaaS backend was convenient.
For GCC and MENA operators, this is not an abstract issue. Many GPU estates are built to serve local enterprises, government-linked entities, Arabic-language data, regulated financial workloads, healthcare use cases or sovereign AI programs. The economics depend on utilization and revenue per megawatt. The governance depends on keeping data, telemetry and operational control inside the required jurisdiction or air-gapped environment. A tool that only works by exporting sensitive traces to an external cloud may not fit the operating model.
Where the control plane has to sit
The useful pattern is to treat AI observability and AI security as tenants of the infrastructure platform, not exceptions to it. ClastIQ is designed for operators that own the hardware and need to run mixed GPU and CPU estates without becoming a hyperscaler. The base requirement is not only Kubernetes. It is a layered control plane across bare metal, VMs, containers, accelerators, storage, identity, quotas and metering.
Bare-metal provisioning matters because the operator needs repeatable builds. MAAS and Juju provide the foundation for commissioning machines, placing services and managing day-two operations. Kubernetes matters because most inference, agent and observability stacks now expect it. VMs matter because some tenants still need appliance-style isolation, legacy images, Windows tooling, or their own kernel and driver assumptions. Ceph matters because checkpoints, embeddings, logs, traces and tenant datasets need block, file or object storage with quota and locality. GPU partitioning matters because an agent workload may need only a fraction of an accelerator while a training job needs a full device.
The platform also needs to meter at the right level. It is not enough to know that node gpu-07 was busy. The operator needs tenant, namespace, VM, job, accelerator profile, storage volume and time interval. When observability tools emit OpenTelemetry traces, those traces should carry tenant labels and workload identity. When an AI security layer blocks an agent action, that event should be attributable to a tenant and policy, not just a hostname.
| Operator problem | Infrastructure control that reduces it |
|---|---|
| Agent traces contain sensitive prompts or code | Tenant-scoped telemetry pipelines, local retention, export policy |
| Small jobs waste whole GPUs | MIG where supported, time-slicing, quota by accelerator profile |
| Teams dispute cost allocation | Per-tenant GPU-hours, CPU-hours, storage and network metering |
| VM and Kubernetes tenants share hardware | Common identity, policy and chargeback model across both |
| Sovereign workload cannot export logs | Air-gapped deployment, local observability, controlled egress |
| Power is constrained before space is | Scheduling policy tied to utilization, quotas and power envelopes |
This is why the HiddenLayer story and the newer observability and agent-security tools should not be read only as software news. They are signals that the GPU estate now needs a governance fabric strong enough for agents, not just batch jobs.
Worked example: why chargeback depends on utilization
Consider an operator with 32 accelerators across four 8-GPU servers. The estate has a theoretical weekly capacity of:
32 GPUs x 168 hours = 5,376 GPU-hours per week.
Before shared orchestration, the operator statically allocates 16 GPUs to an internal platform team, 8 GPUs to a research tenant and 8 GPUs to an application team. Average useful utilization is 35 percent, 20 percent and 15 percent respectively.
The useful GPU-hours are:
- Platform team: 16 x 168 x 0.35 = 940.8 GPU-hours
- Research tenant: 8 x 168 x 0.20 = 268.8 GPU-hours
- Application team: 8 x 168 x 0.15 = 201.6 GPU-hours
- Total useful work: 1,411.2 GPU-hours per week
That is 26.25 percent useful utilization of the estate. The operator still powers, cools, patches and finances all 32 GPUs.
Now assume the same hardware is pooled under quotas. Full-GPU training jobs get reservations. Inference gets priority classes. Agent and notebook workloads use MIG profiles where the accelerator supports it, or time-slicing where that is the right trade-off. Idle quota can be borrowed by lower-priority queues. The average useful utilization rises to 62 percent.
The useful work becomes:
5,376 x 0.62 = 3,333.1 GPU-hours per week.
The improvement is 1,921.9 additional useful GPU-hours per week from the same hardware and power envelope.
Now look at monthly chargeback. Suppose the 32-GPU estate costs USD 960,000 in hardware and is amortised over 36 months, or USD 26,667 per month. Average IT load is 22 kW, facility PUE is 1.5, electricity is USD 0.10 per kWh, and there is USD 4,000 per month in support, networking and spares allocation.
Monthly power cost is:
22 kW x 730 hours x 1.5 x USD 0.10 = USD 2,409.
Total monthly cost is:
USD 26,667 + USD 2,409 + USD 4,000 = USD 33,076.
At 26.25 percent utilization, monthly useful GPU-hours are:
32 x 730 x 0.2625 = 6,132 GPU-hours.
Cost per useful GPU-hour is:
USD 33,076 / 6,132 = USD 5.39.
At 62 percent utilization, monthly useful GPU-hours are:
32 x 730 x 0.62 = 14,483 GPU-hours.
Cost per useful GPU-hour is:
USD 33,076 / 14,483 = USD 2.28.
If the operator charges internal or external tenants USD 3.20 per useful GPU-hour, monthly revenue is USD 46,346 at the improved utilization. With facility-adjusted load of 33 kW, this is about USD 1.4 million per MW-month, or about USD 16.9 million per MW-year. The exact tariff will vary by country, hardware generation, financing and support model. The point is simpler: observability and security controls only become economically useful when they are tied to allocation, metering and quota.
Without that link, the operator may know that an agent did something risky, but not which tenant should pay for the capacity it consumed or which quota should be reduced.
A practical implementation pattern
For a small-to-mid GPU operator, the implementation sequence should be boring and enforceable.
First, create the tenant model before onboarding observability tools. A tenant should have an identity boundary, namespaces or projects, VM ownership, storage quotas, accelerator quotas, cost center labels and data classification. This is not paperwork. It is the metadata that lets telemetry, policy and billing join together later.
Second, standardize the accelerator profiles. Full GPU, MIG slice, time-sliced GPU, CPU-only high-memory node and storage-heavy node should be treated as schedulable products. NVIDIA, AMD and mixed estates should fit the same operator model even if the device plugins and partitioning features differ. Hardware-agnostic does not mean pretending all GPUs are identical. It means presenting clean quota and metering primitives above the hardware details.
Third, run observability locally unless policy explicitly allows export. OpenTelemetry is useful because it gives a common language for traces and metrics across agents, APIs and infrastructure. But the collector, storage backend and retention policy must be tenant-aware. Prompt logs and tool-call traces may contain code, personal data, credentials, customer records or government material.
Fourth, put the AI security layer on the same identity spine as the rest of the estate. If a coding agent proposes a shell command, opens a repository, calls a secrets API or reaches an inference endpoint, the event should carry tenant and workload identity. Security alerts that cannot be mapped to owner, quota and policy become manual tickets.
A representative namespace bootstrap might look like this. Resource names will vary by accelerator vendor and Kubernetes device plugin, but the intent is stable:
kubectl create namespace tenant-finance-ai
kubectl label namespace tenant-finance-ai tenant=finance-ai chargeback=cc-410 data-class=restricted
kubectl annotate namespace tenant-finance-ai telemetry.retention=30d egress.policy=local-only
cat <<'EOF' | kubectl apply -f -
apiVersion: v1
kind: ResourceQuota
metadata:
name: finance-ai-quota
namespace: tenant-finance-ai
spec:
hard:
requests.cpu: '256'
requests.memory: 1Ti
requests.storage: 20Ti
requests.nvidia.com/gpu: '4'
EOF
The same idea should apply to VMs and bare-metal reservations. A tenant consuming a VM with two full GPUs should appear in the same metering report as a tenant consuming four MIG slices in Kubernetes. Storage should be metered from Ceph pools or volumes. Chargeback should not depend on whether the workload entered through a VM portal, a Kubernetes namespace, a batch queue or a managed notebook.
Partitioning choices for agents and LLM workloads
Agent workloads are often bursty. A coding agent may run many CPU-heavy steps, then call an LLM endpoint, then run tests, then sit idle. A document agent may need high memory and storage more than full GPU time. A fine-tuning job may need continuous full-device access for hours. Treating all of them as identical GPU jobs creates both security and utilization problems.
MIG, where supported, is useful when the operator needs hard GPU partitioning with predictable slices. It can keep small inference or agent-support workloads from occupying a whole accelerator. Time-slicing is useful when the goal is statistical sharing and the workloads can tolerate contention. Full GPU allocation remains necessary for many training, fine-tuning and latency-sensitive inference jobs. CPU-only and high-memory pools are also important because not every AI control-plane component belongs on a GPU node.
A ClastIQ-style operating model keeps these choices visible to quota and metering. A tenant may receive two full-GPU reservations for training, eight small GPU slices for development, a CPU quota for agents and a Ceph object bucket for traces. Another tenant may be denied internet egress and required to use only local model endpoints. A third may be allowed to export aggregate metrics but not raw prompts.
This is where policy engineering becomes practical rather than theoretical. The operator can say: restricted tenants may run agents only from approved images; telemetry must stay local; GPU slices are capped; storage retention is 30 days; full GPUs require reservation; preemptible queues may borrow unused quota overnight. Those policies reduce ambiguity during incidents and during billing disputes.
Storage, telemetry and sovereign operation
AI observability increases storage pressure. Traces, logs, prompts, model responses, embeddings, vector indexes, checkpoints and audit events all have different retention and sensitivity. If every team runs its own datastore, the operator loses quota control and backup discipline. If everything is dumped into one shared bucket, the operator loses tenant isolation.
Ceph gives an owned-hardware operator a way to provide block, file and object storage from the same estate while keeping pools, quotas and replication policies explicit. The important design decision is to align storage classes with tenant policy. Short-lived debug traces do not need the same retention as audit records. Model checkpoints may need high throughput but not indefinite retention. Security events may need immutability or restricted administrator access.
For GCC and MENA estates, the sovereign angle is often decisive. Local organizations may require that prompts, source code, telemetry and model artifacts remain inside a country, campus or air-gapped site. Air-gap support is not only about installing without internet access. It also means mirrored packages, local registries, offline updates, internal certificate authorities, local observability backends and support procedures that do not require shipping sensitive logs to a foreign SaaS system.
Power is the other local constraint. In dense GPU rooms, the limiting factor may be available kW, cooling or transformer capacity rather than rack units. A platform that meters GPU-hours but ignores power can still produce poor economics. Operators should track revenue, internal recovery or research output against facility-adjusted power. Revenue per MW is a useful discipline because it links scheduling policy to the physical constraint that actually limits expansion.
What ClastIQ contributes to this layer
ClastIQ does not replace specialized AI observability or security products. The better pattern is to give those products a governed substrate. ClastIQ brings the operator controls around the fleet: bare-metal provisioning with MAAS and Juju, Kubernetes and VM tenancy on the same hardware, GPU partitioning through MIG or time-slicing where appropriate, Ceph-backed storage, quotas, per-tenant metering, chargeback, air-gap deployment patterns and local human support.
That matters because the agent security layer needs facts from the infrastructure: tenant identity, workload identity, resource limits, egress policy, storage location and accelerator allocation. The observability layer needs the same facts to avoid becoming an ungoverned data exhaust. Finance needs them to recover cost. Facilities needs them to understand which workloads justify power. Security needs them to respond without guessing who owns a process.
For an operator with a few racks, the objective is not to copy every hyperscaler service. It is to make the estate governable enough that more tenants can safely share it. That means higher utilization without losing isolation, better chargeback without spreadsheets, and local control over data and telemetry.
What to do this week
- Inventory every AI workload that uses GPUs, CPUs, agents, observability stores or shared model endpoints, and assign a tenant owner and cost center.
- Define three to five accelerator profiles, such as full GPU, small slice, time-sliced development, CPU-only agent and high-memory CPU, then map quotas to them.
- Decide which telemetry may leave the site, which must stay local, and which must be redacted before storage or export.
- Run a one-week utilization and power baseline: GPU-hours allocated, GPU-hours useful, storage growth, average kW and top idle reservations.
- Create a chargeback model using real estate cost, power cost and support cost, then publish the cost per useful GPU-hour at current and target utilization.
- Test one tenant onboarding path end to end across namespace or VM creation, storage quota, GPU quota, telemetry labels, egress policy and metering report.
ClastIQ runs this control plane on the operator’s own hardware: request a demo.
