Flow-Control-Paused spikes

654 views
Skip to first unread message

Aleksey Sanin

unread,
Jul 26, 2013, 4:42:09 PM7/26/13
to codersh...@googlegroups.com
Hello,

We run Galera cluster on 3 nodes with one write node and two read nodes
(including one master for regular replication). Once or twice a day we
spikes for the Flow-Control-Paused on one (usually write node) or all
the nodes (see attached screenshot). This causes some of the writes to
time out and fail.

Would appreciate any ideas or suggestions on the cause or ways to debug
it. The DB nodes don't have any other processes running and backups are
done at night (and we don't usually see any spikes during the backup).

Thank you in advance,

Aleksey

Screen Shot 2013-07-26 at 1.38.21 PM.png
signature.asc

Alex Yurchenko

unread,
Jul 27, 2013, 9:38:08 AM7/27/13
to codersh...@googlegroups.com
Hi,

wsrep_flow_control_paused is a global cluster status, normally it should
be the same on all nodes. If it is not - it is either some obscure bug
or sampling error.

what you should do when you see that is take node of
wsrep_flow_control_sent - this will show which node is in trouble.

so far the only likely reason for replication pause is a stuck node,
either due to some IO-intensive task or some OS issues (like huge pages
defragmentation, make sure you have huge pages disabled and IO scheduler
is not cfq)

the most likely reason is a huge transaction (especially if it is a
delete on a table without primary key. Then it is bound to be very long
running on slaves and no amount of parallel applying will help you
there).

Regards,
Alex
--
Alexey Yurchenko,
Codership Oy, www.codership.com
Skype: alexey.yurchenko, Phone: +358-400-516-011

Aleksey Sanin

unread,
Jul 27, 2013, 2:34:47 PM7/27/13
to Alex Yurchenko, codersh...@googlegroups.com
Hi Alex,

Thanks for the reply. We do have all the stats dumped and
wsrep_flow_control_sent is always 0. We also dumped IO stats
and there is nothing unusual there either. The DB itself has
no tables w/o primary key and there are no deletes either.

The two OS settings are set correctly (huge pages disabled,
IO scheduler set to deadline).

Any ideas of how to debug this?

Aleksey
signature.asc

Alex Yurchenko

unread,
Jul 27, 2013, 5:01:25 PM7/27/13
to Aleksey Sanin, codersh...@googlegroups.com
On 2013-07-27 21:34, Aleksey Sanin wrote:
> Hi Alex,
>
> Thanks for the reply. We do have all the stats dumped and
> wsrep_flow_control_sent is always 0.

You must have a bug in monitoring. flow control is maintained by sending
special beacons, so you simply can't have a pause without somebody
sending such beacon.

One problem you may face there is that currently every time SHOW STATUS
command is issued, it clears the counters. If you have more than one
monitoring tool, then they can steal data from each other.

Do you monitor wsrep_local_recv_queue? What does it show on master and
slaves when this thing happens?

Aleksey Sanin

unread,
Jul 27, 2013, 5:10:56 PM7/27/13
to Alex Yurchenko, codersh...@googlegroups.com
One tool, runs every second through cron. wsrep_local_recv_queues is
also zero (there are absolutely no changes in all the other stats when
the spike happens).

Just for the record, we run 23.2.6

Aleksey
signature.asc

Aleksey Sanin

unread,
Aug 8, 2013, 1:54:22 AM8/8/13
to codersh...@googlegroups.com
We just had another spike that caused brief production downtime.
Again, all other graphs are at 0 while the Flow-Control-Paused
on all 3 servers goes up.

Any ideas on the cause or ways to debug it are greatly appreciated.

Thanks

Aleksey





On 7/27/13 2:01 PM, Alex Yurchenko wrote:
signature.asc

Alex Yurchenko

unread,
Aug 9, 2013, 5:46:25 PM8/9/13
to codersh...@googlegroups.com
On 2013-08-08 08:54, Aleksey Sanin wrote:
> We just had another spike that caused brief production downtime.
> Again, all other graphs are at 0 while the Flow-Control-Paused
> on all 3 servers goes up.

There must be a bug in monitoring software that you use. Maybe a bug in
graph plotting. Have you looked at raw numbers?

Enabling general query log could have helped with tracking the offending
query, but it is terribly expensive.

Dumping

SHOW FULL PROCESSLIST;
SHOW GLOBAL VARIABLES\G
SHOW STATUS LIKE 'wsrep%';
SHOW ENGINE InnoDB STATUS\G

during such stall can also be helpful.

Reagards,
Alex

Aleksey Sanin

unread,
Aug 9, 2013, 5:50:48 PM8/9/13
to Alex Yurchenko, codersh...@googlegroups.com
Thanks for the reply. I actually did more searching and I wonder if this
bug might be related:

https://bugs.launchpad.net/galera/+bug/1180792

Aleksey
signature.asc

Alex Yurchenko

unread,
Aug 9, 2013, 7:31:35 PM8/9/13
to Aleksey Sanin, codersh...@googlegroups.com
On 2013-08-10 00:50, Aleksey Sanin wrote:
> Thanks for the reply. I actually did more searching and I wonder if
> this
> bug might be related:
>
> https://bugs.launchpad.net/galera/+bug/1180792
>
> Aleksey

That's what I suggested in the very beginning, but it is just a matter
of data collection. If you have ONLY one client that polls status
variables, you WILL see non-zero wsrep_flow_control_sent/recv before the
pause. The question is - can you see a value of 1 on the chart?

Aleksey Sanin

unread,
Aug 9, 2013, 9:00:57 PM8/9/13
to Alex Yurchenko, codersh...@googlegroups.com
Yep, makes sense. Unfortunately I don't see it. It's one client per node
in the crontab (runs every minute).


Aleksey
signature.asc
Reply all
Reply to author
Forward
0 new messages