Hi,
wsrep_flow_control_paused is a global cluster status, normally it should
be the same on all nodes. If it is not - it is either some obscure bug
or sampling error.
what you should do when you see that is take node of
wsrep_flow_control_sent - this will show which node is in trouble.
so far the only likely reason for replication pause is a stuck node,
either due to some IO-intensive task or some OS issues (like huge pages
defragmentation, make sure you have huge pages disabled and IO scheduler
is not cfq)
the most likely reason is a huge transaction (especially if it is a
delete on a table without primary key. Then it is bound to be very long
running on slaves and no amount of parallel applying will help you
there).
Regards,
Alex
--
Alexey Yurchenko,
Codership Oy,
www.codership.com
Skype: alexey.yurchenko, Phone: +358-400-516-011