This release makes it possible to merge traces from different machines onto one timeline, introduces Memscope and Memory Overview for end-to-end memory investigations, lets you keep a parsed trace running and analyze it through lightweight remote clients, adds Zstd trace compression, and substantially improves Perfetto’s AI-assisted analysis workflows. It also introduces a common stack-sampling format, new export options, and a collection of UI and SDK improvements.
Perfetto can now open several traces together and merge their contents onto a single timeline! This makes it much easier to investigate behavior spanning multiple VMs, machines, or separately recorded trace sources without manually converting them into one trace first.
In the UI, select Open multiple trace files, select several files in the normal file picker, or drop multiple files onto the page. A new dialog lets you configure how the traces should be merged, shows whether any events would be dropped, and can export the result as a self-contained .tar archive.
A powerful use of this feature is to visualize traces that span machines and VMs. For example, to visualize traces across a phone and a backend service the phone is communicating with:
Use the new util merge command which helps avoid common mistakes when merging traces. The resulting trace can be opened in the UI or in the trace processor:
# Merge traces from the same machine recorded with consistent clocks. For multi-machine or multi-clock setups, use the JSON # manifest as described below. trace_processor util merge \ -o merged.tar \ perf.data \ perfetto.pftrace # Now you can opened merged.tar in the UI *or* load it in trace processor: trace_processor merged.tar
For more complex setups, a perfetto_manifest JSON file can assign machine names, relate clock domains, and apply fixed clock offsets. The UI can generate both the archive and manifest interactively, providing a convenient starting point for scripted or repeatable workflows.
See the merging traces documentation for much information examples of how to merge traces in common setups, details of the manifest format and details on how to generate pre-merged traces for loading in the UI, without using trace_processor merge
Perfetto v58 introduces a connected memory-analysis workflow spanning live monitoring, trace recording, initial triage, and detailed investigation.
Memscope is a live memory profiler for Android and Linux. It guides you through connecting to a device or host, displays system-wide and per-process memory statistics as they change, and provides shortcuts for recording a memory trace for a selected process. Instead of constructing a memory-focused recording configuration by hand, you can move directly from observing suspicious memory behavior to capturing the information needed to explain it.
After recording, Memory Overview provides a high-level breakdown of where a process’s memory is being used. It combines information from smaps snapshots, ART heap dumps, and native heap profiling into one starting point, with links into more specialized tools such as Heap Dump Explorer.
Memory Overview also opens automatically when a trace contains smaps snapshots, making memory traces immediately actionable rather than dropping users onto a generic timeline.
Together, Memscope and Memory Overview provide a path from “this process appears to be using too much memory” to the appropriate native, managed, or system-level investigation.
See the Memscope and Memory Overview guide for a complete recording and analysis walkthrough.
Remote Trace Processor analysis is not new: trace_processor server http already allowed the Perfetto UI to use a native Trace Processor instance instead of the browser’s WebAssembly implementation.
The command-line experience, however, had significant gaps. The Trace Processor CLI could not connect to an existing server, and running several server instances meant manually assigning port numbers, remembering which trace was attached to each port, and dealing with port clashes.
Perfetto v58 turns this into a complete client-server workflow. A trace can be loaded once into a properly named background session, and any number of lightweight command-line clients can connect to it using --remote:
# Parse the trace once and keep it in a named session. trace_processor server unix \ --name mysession \ --daemonize trace.pftrace # Run queries against the existing session. trace_processor query \ --remote mysession \ "SELECT count(*) FROM slice"
Named local sessions use Unix sockets, so there is no port to allocate or remember. The session name identifies both the running Trace Processor and the trace it contains. The new remote client can also connect to a trace_processor server http instance using --remote host:port, bringing the existing HTTP/WebSocket server workflow to the command-line interface.
Session state persists across invocations: tables created by one query and modules included by one command remain available to the next, allowing expensive intermediate results to be materialized once and reused. The result is a much more usable workflow for large traces: parse once, give the session a meaningful name, and analyze it repeatedly without juggling ports or restarting Trace Processor.
See the command-line analysis guide for session management, remote addresses, and idle-timeout options.
Perfetto now supports Zstd as a trace compression option alongside deflate. Zstd generally produces smaller traces than deflate at a similar speed, reducing the storage and transfer costs of longer or more data-intensive recordings.
Compression is selected through the new TraceConfig.compression field:
compression {
zstd {}
}The compression level can also be controlled explicitly:
compression {
zstd {
level: 9
}
}Level 0, or leaving the level unset, uses Zstd’s default level of 3, which provides a good balance for most traces. The older compression_type field remains honored for compatibility but is now deprecated in favor of TraceConfig.compression.
See the trace configuration documentation for level selection, compatibility behavior, and SDK build requirements.
Perfetto’s installable AI skill can now keep a trace loaded across an investigation. Instead of reparsing the same trace for every question, the skill opens a persistent Trace Processor session and reuses it for subsequent queries. This makes iterative analysis significantly faster and allows SQL state and materialized intermediate results to carry across steps.
The skill also gains:
TraceConfig reference and exemplar configurations that an agent can adapt when a question requires a custom recording.perfetto-ai-skill.zip asset in each GitHub release, allowing the skill to be installed on machines that cannot reach GitHub at installation time.See the Using AI with Perfetto guide for installation instructions and example investigations.
Perfetto now defines a common stack-sampling format that any profiler can emit.
The new TracePacket.stack_sample packet attributes a callstack to a thread, process, or stackful asynchronous context such as a goroutine or fiber. Samples are measured against an explicit counter timebase—such as wall time, CPU cycles, or allocated bytes—and can carry additional follower counters.
Callstacks can be fully inline or use Perfetto’s interning mechanism. Per-sequence StackSampleDefaults identify the profiler and declare common counters once for the whole sampling stream.
Trace Processor exposes the results through a common set of tables:
stack_samplestack_sample_sessionstack_sample_task_contextstack_sample_execution_contextstack_sample_async_contextstack_sample_counter_trackstack_sample_counterThese tables provide one profiler-neutral model while continuing to share the existing stack_profile_* callstack tables. Linux perf’s specialized perf_sample, perf_session, and perf_counter_track tables remain available for perf-specific information, including counter-only samples.
This gives profiler authors a standard route into Perfetto’s SQL and visualization capabilities without having to adopt another profiler’s transport format.
Perfetto v58 adds two Trace Processor export formats for workflows that need to reuse or move already-parsed data.
# Reloadable by the same Trace Processor version. trace_processor export perfetto \ -o parsed.tar \ trace.pftrace # Standard Arrow files for external analysis tools. trace_processor export arrow_tar \ -o tables.tar \ trace.pftrace
The perfetto format stores parsed data in an archive that another Trace Processor instance of the same version can load. This avoids repeatedly parsing a large source trace when a persistent server session is not appropriate.
The arrow_tar format writes the built-in tables as standard Apache Arrow files inside a TAR archive. The format is stable across Perfetto versions and can be consumed by tools such as pandas, Polars, and PyArrow. It is intended for interoperability and cannot be loaded back into Trace Processor.
Together with client-server mode, these formats provide three distinct options: keep a live parsed session, save a reloadable parsed archive, or export tables into a standard analytics ecosystem.
Flamegraphs now have clearer direction and filter controls, frame highlighting, and case-insensitive regular-expression filtering.
Callstack sample tracks and flamegraphs also share a common implementation over the new stack_sample table. Tracks and area-selection tabs remain separate for each profiler source, preserving source-specific workflows, while any source that records counters can now support counter-weighted flamegraphs.
The Perfetto homepage now lists recently opened traces, making it easier to return to an investigation without locating the file again.
The slice aggregation table can now pivot on values stored in event arguments. This makes it possible to compare groups based on dimensions encoded in args without first reshaping the data through a custom SQL query.
Additional UI improvements include:
F1 on Firefox.More than one tracing session can be active on a device at the same time—for example, a developer recording manually while a system service or automated test is also tracing. Until now, the resulting trace did not show when those other sessions started or stopped, making interference between recordings difficult to explain.
Perfetto can now record the lifecycle of every other session that overlaps with yours. Enable it in the recording config:
builtin_data_sources {
enable_concurrent_session_events: true
}The trace records when each concurrent session is configured, started, begins shutting down, or becomes disabled. Trace Processor and the UI turn those events into one state track per session under Concurrent tracing sessions, so the overlap is visible alongside the activity being investigated.
The process_stats data source now writes an explicit Thread message for every process’s main thread instead of leaving it implied by the process entry. Traces recorded from procfs alone therefore retain main-thread names in Trace Processor and the UI.
Trace-level notes have been generalized into attributes. TraceConfig.trace_attributes and perfetto --add-attribute replace TraceConfig.notes and --add-note, allowing a recording to carry structured context such as its build identifier, experiment name, test case, or collection environment. These attributes are available from Trace Processor without requiring a timestamped event or metadata track.
The C SDK’s shared-library configuration has been simplified:
PERFETTO_SDK_SHLIB_IMPLEMENTATION when building the SDK into a shared library.PERFETTO_SDK_SHLIB when consuming it as a shared library.The C data-source ABI also gains PerfettoDsGetTimestamp() and PerfettoDsGetDefaultClockId(). These APIs return timestamps using Perfetto’s preferred trace clock, removing the need for callers to select an operating-system-specific clock themselves. The Rust SDK provides the equivalent through DataSourceTimestamp::now().
On Apple platforms, dev.perfetto.clock_sync signpost events now fire on iOS as well as macOS, allowing Instruments traces captured on iOS to be synchronized with Perfetto traces.
The SQL regexp() function now accepts an optional third argument for matching flags: i selects case-insensitive matching and c explicitly selects case-sensitive matching.
This release also fixes:
CpuInfo packet.Perfetto v58 includes several API and tooling changes worth checking when upgrading:
TrackEvent.Callstack has moved to the top-level InlineCallstack message in trace/profiling/inline_callstack.proto. The wire format is unchanged, but code referring to the generated TrackEvent.Callstack or TrackEvent::Callstack types must move to InlineCallstack.TraceConfig.notes and perfetto --add-note have been replaced by TraceConfig.trace_attributes and --add-attribute.TraceConfig.compression supersedes the deflate-only compression_type field.PERFETTO_SHLIB_SDK_IMPLEMENTATION has been renamed to PERFETTO_SDK_SHLIB_IMPLEMENTATION, and PERFETTO_SDK_DISABLE_SHLIB_EXPORT has been removed.traceconv: the standalone tool is deprecated in favor of equivalent trace_processor subcommands. The traceconv wrapper continues to fetch Trace Processor and run it under the old name, so existing invocations continue to work for now.--allow-sql-file-access.EXPORT_JSON: the integer file-descriptor form has been removed. SQL file writes no longer require --dev.Happy Tracing!
For contributor acknowledgements, downloads, assets, and the complete release record, see Perfetto v58.2 on GitHub.