The Why

Why IT Observability Matters

Five reasons observability matters in increasingly complex IT environments.

01

Increasing Service Complexity

Gain unified visibility across the entire IT environment
Understand the scope of service impact
Break down data silos
Achieve end-to-end visibility
02

Delayed Root Cause Analysis

Reduce root cause analysis time
Correlate related events
Lower MTTR
Accelerate incident response
03

Resource Optimization

Analyze CPU and memory utilization
Identify over- and under-allocated resources
Support capacity planning
Reduce infrastructure costs
04

Operational Inefficiency

Unify operational data in a single dashboard
Automate events and alerts
Reduce repetitive manual tasks
Improve operational productivity
05

Difficulty Assessing Service Impact

Improve service availability
Detect anomalies early
Prevent service disruptions
Maintain a stable operating environment
Anomaly Detection

Anomaly Detection Is Just the Beginning.
What Matters Is How You Use It.

Make Kubernetes operations faster and smarter with onTune’s anomaly detection widget.

{{ t.num }}
{{ t.label }}
{{ t.desc }}
{{ t.alt }}

Anomaly Detection, Reimagined with onTune

Go beyond anomaly alerts by seeing exactly how far each metric deviates from its baseline.

Scope
4Areas

Monitor anomalies across four key areas: nodes, workloads, placement/capacity, and the control plane.

Granularity
10Levels

Measure the degree of deviation from the baseline across 10 levels for more precise anomaly assessment.

Interval
2Seconds

Update anomaly status at intervals as short as two seconds for near-real-time detection.

Window
10Minutes

Review status changes over the last 10 minutes at a glance.

See Anomalies at a Glance.
Dive Deeper in One Click.

View anomaly severity and detection criteria at a glance, then drill down into the details with a single click. Color-coded indicators and bars make results easy to understand.

Tailored Baselines
for Every Cluster

Set cluster-specific baselines based on workload characteristics for more accurate anomaly detection.

Beyond Monitoring

See the Entire Service. Understand the Root Cause.

onTune Observability brings metrics, logs, traces, events, and topology into one unified context—helping you move from knowing what failed to understanding why.

{{ s.icon }} {{ s.name }}
User, App, Kubernetes, Server, DB, Network, Cloud — all sources converge into onTune Observability.
UNIFIED CONTEXT
onTune Observability
{{ g.name }} {{ g.desc }}
{{ flowDetail }}

Metrics — Collect performance metrics across systems and services. Identify when performance degradation begins by comparing changes across key metrics and pinpoint where the issue first emerged.

Logs — Correlate application and system logs. Correlate related logs to pinpoint the cause of an issue and analyze error messages alongside system status.

Traces — Trace request flows across distributed environments. Trace slow requests across services to identify bottlenecks and understand how performance issues propagate through the call flow.

Events — Track key changes, alerts, and system events. See what changed before and after an incident. Correlate deployments, configuration changes, and failure events to quickly narrow down potential causes.

Topology — Visualize relationships across systems and services. Visualize relationships across applications, pods, servers, databases, and networks to understand service dependencies and the impact of failures.

onTune BENEFITS
troubleshoot
Identify the Root Cause
Pinpoint where the issue started and why it occurred.
share
Assess the Impact
Identify which services and users are affected by the incident.
shield
Prevent Recurrence
Analyze failure history and patterns to reduce the risk of recurring issues.
The Platform

All Your IT Environments. One Platform.

From infrastructure to applications, select the area you need and take a closer look.

Kubernetes Monitoring, Public Cloud Monitoring, GPU Monitoring, Network Performance Monitoring, Server Monitoring, VM/Xen Monitoring, Application Monitoring, Network Monitoring, Database Monitoring, URL Monitoring
{{ selected.icon }} {{ selected.name }}
{{ selected.feature.placeholder }}

{{ selected.feature.title }}

check_circle {{ b }}

Kubernetes Monitoring — Observation of container environment from Cluster to Pod

Real-Time Visibility into Your Kubernetes Environment

  • Monitor cluster, node, pod, and namespace status from a single view
  • Understand overall operational status at a glance with a unified summary dashboard

Gain Multi-Layer Visibility Across Kubernetes

  • Analyze Kubernetes data across layers, from clusters to workloads
  • Maintain consistent visibility as your Kubernetes environment scales
  • Analyze the performance and health of individual resources

Turn Collected Data into Deeper Operational Insights

  • Identify root causes when issues occur
  • Analyze performance data and event history
  • Track configuration changes across Kubernetes objects
  • Detect operational patterns and anomalies

Public Cloud Monitoring — Monitoring cloud resources such as AWS, Azure, NCP, etc.

Unified Monitoring Across Public Cloud Environments

  • Monitor multiple cloud environments from a single view
  • Collect data through agents and native cloud management services across supported public clouds

Monitor Key AWS and Azure Services

  • Collect data from AWS services including EC2, S3, EBS, ELB, RDS, and EKS, along with billing information
  • Collect data from Azure services including Virtual Machines, Blob Storage, Load Balancer, SQL Database, and AKS, along with billing information

Flexible Reporting and Historical Analysis

  • Quickly create the reports you need from the Report menu
  • View collected data at 1-, 10-, 30-, or 60-minute intervals
  • Compare the same time period across yesterday, last week, and last month to identify trends

GPU Monitoring — Track GPU utilization and performance of AI infrastructure

Unified Monitoring Across GPU Environments

  • Monitor GPU environments across Linux, Windows, and Kubernetes from a single view
  • Track GPU utilization in real time down to the process (PID) level in Linux and the pod level in Kubernetes

Precise Monitoring of Physical GPUs and Time-Sliced Resources

  • Track asset information and hardware health for all physical GPUs in real time
  • Monitor time-sliced resource allocation and kernel-level ECC errors in one view

Precise Monitoring for Logical GPUs (MIG)

  • Track MIG instance assets and resource allocation status in real time
  • Track 1:1 mappings between MIG instances and Kubernetes pods
  • Quickly identify idle, underutilized, and unallocated GPU resources

Network Performance Monitoring — Real-time analysis of communication quality and TCP performance between hosts

Real-Time Network Health at a Glance

  • Monitor network health in real time with host status and health scores
  • Detect bottlenecks using RTT, jitter, packet loss, retransmissions, and throughput
  • Find network failure causes by correlating connection, kernel, and event data

Visualize Network Relationships with Topology

  • Automatically map connections across hosts, databases, and external systems
  • Monitor latency, jitter, throughput, sessions, and connections by node
  • Quickly locate failures and affected areas through visual alerts and events

Pinpoint Network Performance Issues

  • Monitor TCP throughput, RTT, jitter, retransmissions, and packet loss in real time
  • Compare host and traffic trends to identify latency, loss, and retransmission issues
  • Analyze connection rates and TCP states to find the causes of performance degradation

Server Monitoring — Analyze server performance, configuration, and events in seconds

Flexible Server Analysis Views

VM/Xen Monitoring — Visibility of virtualization infrastructure, including VMware and XenServer

Unified VMware Host Monitoring

  • Monitor and analyze VMware environments running on x86 infrastructure
  • View overall host group status and drill down into individual hosts through Overview, Group, and Single views

Comprehensive VMware Host Performance Metrics

  • Monitor 30+ performance metrics across VMware hosts
  • Track CPU, memory, resource pools, network, disk, datastores, system information, and events in one dashboard

Comprehensive XenServer Performance Monitoring

  • Monitor 20+ performance metrics across your XenServer environment
  • Analyze performance, logs, datastores, ping, and port scan data from one interface

Application Monitoring — Transaction, user experience, service call flow analysis

Real-Time APM Performance Monitoring

  • Monitor key APM metrics such as TPS, response time, error rate, and active transactions
  • Drill down from service to instance to transaction for JVM, heap, GC, thread, and resource analysis
  • Identify failures and performance degradation through slow transactions, errors, and resource usage

See Performance from the User’s Perspective

  • Monitor real user experience in real time based on actual sessions and page performance
  • Identify slow pages and performance issues through Core Web Vitals, page load waterfalls, and device/browser data
  • Trace user experience issues from the frontend through APIs, services, and traces

Visualize End-to-End Service Flows

  • Visualize call relationships from clients and web tiers to WAS, external services, and databases
  • Identify bottlenecks and failure points using response time, throughput, errors, and call performance
  • Trace individual transaction paths to understand service call flows and pinpoint root causes

Network Monitoring — Network device status and traffic visibility

Extend Monitoring with Private OIDs

  • Use standard SNMP public OIDs and extend monitoring with vendor-specific private OIDs
  • Collect custom performance metrics and generate reports without additional vendor-specific development

Integrated Multi-Vendor Monitoring

  • Monitor devices from multiple vendors in one view and collect key performance metrics centrally
  • Track configuration data over time and compare changes on a daily basis

Customize Your Monitoring View

  • Save frequently used monitoring layouts and quickly restore them whenever needed
  • Add, remove, and arrange widgets to match your monitoring needs

URL Monitoring — URL health monitoring and performance analysis

Monitor URL Health at a Glance

  • Monitor URL status and response times at a glance
  • Quickly identify status changes with visual indicators
  • Track key URL events in real time

Manage URL Groups and Event Policies

  • Group multiple URLs for easier management
  • Define events based on status, response time, SSL, and other conditions
  • Configure alert targets, conditions, and recipients for faster response

Database Monitoring — In-depth analysis of DB performance such as Oracle, MySQL, etc.

Agentless Database Monitoring

  • Collect database data through queries without installing agents on DB servers
  • Support major databases including Oracle, SQL Server, DB2, PostgreSQL, and MySQL
  • Configure alert targets, conditions, and recipients for faster incident response

Detect Database Issues with Custom Events

  • Configure hardware fault and health-check events
  • Define event conditions based on operational requirements
  • Detect process, log, and other database-related events
The Flow

From Detection to Resolution, All in One onTune

01

Detect

Detect anomalies in real time and capture the data that matters.

02

Correlate

Correlate metrics, logs, and traces to understand the scope of impact.

03

Analyze

Identify performance bottlenecks and pinpoint root causes faster.

04

Resolve

Resolve issues, validate recovery, and prevent recurrence.

Solve Critical IT Operations Issues Faster

trending_down

Poor Service Performance

Identify the root cause of overall service performance degradation.

schedule

API Performance Delays

Identify the causes of increased API response times.

restart_alt

Kubernetes Pod Restarts

Trace pod restart causes and related events.

database

Slow Database Response

Diagnose slow queries and database performance issues.

network_check

Network Performance Degradation

Detect latency, packet loss, and other network issues.

account_tree

Failure Impact Analysis

Assess how failures affect services and business operations.