Hii,
I have a synchronous gRPC C++ client that streams large telemetry messages to a remote collector using a persistent ClientReaderWriter stream. I am seeing monotonically increasing RSS on the sender while the stream is active and healthy, and I suspect gRPC's transport layer is buffering serialized data faster than the collector consumes it.
SetupgRPC C++ version: 1.50.1
API: synchronous gRPC, not the async completion-queue API
Transport: insecure gRPC over TCP
Stream type: ClientReaderWriter<SubscribeResponse, PublishResponse>
Persistent stream: created once per subscription, reused for all messages (one ClientContext, one stream object)
Message size: ~4–6 MB serialized protobuf per Write()
Write rate: one message every ~10 seconds
Collector read rate: slower than the write rate; each blocking Write() returns after ~6–16 seconds
Channel arguments:
args.SetMaxReceiveMessageSize(25 * 1024 * 1024);
args.SetMaxSendMessageSize(25 * 1024 * 1024);
args.SetInt(GRPC_ARG_ENABLE_CHANNELZ, 0);
args.SetInt(GRPC_ARG_ENABLE_CENSUS, 0);
args.SetInt(GRPC_ARG_MINIMAL_STACK, 1);
Over several hours, the sender process RSS grows by hundreds of MB (e.g., ~942 MB → ~1.5 GB). The growth is roughly proportional to the number of successful Write() calls, and the rate is larger than the raw payload size (on the order of 7–8 MB RSS growth per 4–5 MB payload).
What I have already ruled out
1. Application-level queue leak: I cap the work queue at 20 entries and log its estimated memory; it stays bounded while RSS grows.
2. Serialization leak: the SubscribeResponse object is stack-local and Clear()'d after each Write().
3. Thread leak: thread IDs are stable; no churn.
4. Allocator retention: tcmalloc stats show current_allocated_bytes growing alongside RSS; free caches (pageheap_free_bytes, thread_cache_free_bytes, etc.) are flat. This is real in-use heap growth, not retained-but-free memory.
5. Error-path leak: the growth happens during successful streaming with no Write() failures, no stream closes, and no reconnections.
Code pattern
// Synchronous API: stream created once
_streamContext.reset(new grpc::ClientContext());
_stream = stub->Publish(_streamContext.get());
// Blocking loop on the sender thread
while (running) {
gnmi::SubscribeResponse publish_data = BuildNotification(...);
// BLOCKING Write(): returns after ~6-16 s because collector is slow
bool ok = _stream->Write(publish_data);
publish_data.Clear();
}
Key observation
Write() is synchronous and blocking, so there is only one outstanding message at the application level. Yet RSS still grows unbounded. This suggests gRPC's transport/HTTP/2 layer is retaining serialized data internally even after the synchronous Write() call returns.
Questions
1. Is this unbounded transport buffering expected behavior for a slow reader with a synchronous persistent stream? If so, what is the correct way to limit in-flight bytes or apply backpressure?
2. Does the synchronous gRPC C++ API provide any way to query or cap the number of serialized-but-not-yet-ACK'd messages/bytes in a ClientReaderWriter?
3. Should I use ResourceQuota, HTTP/2 flow-control channel args, or a smaller SetMaxSendMessageSize to bound this? If yes, what are the concrete knobs?
4. Is periodically calling WritesDone() + Finish() and recreating the stream the only practical way to force a flush of transport buffers, or is there a cleaner API?
5. Are there known issues where a large SetMaxSendMessageSize combined with the sync API causes excessive internal buffering?
Any pointers to the right gRPC knobs would be appreciated.