Customer support

Incident Alerts in Customer Support: When One Problem Suddenly Spikes

How to tell when a problem is blowing up in support, how to set the alert threshold and window so it's neither late nor noisy, and what to do after the alert fires.

In this article
  1. A scattered problem vs. a spike
  2. A good alert has three parameters
  3. Where the alert shows up
  4. What to do after the alert
  5. Proactive customer notices
  6. Common mistakes
  7. Frequently asked questions

It's 10 a.m. One agent has received three messages since morning about "verification code isn't arriving," another agent two, and neither knows the other has heard the same thing. By the time the team lead notices at noon, dozens of people have written in, the queue is long, and each customer has separately received the same "we're looking into it" reply.

An incident is usually recognized this way: not by a single message, but by acceleration. The sooner you see that a problem is suddenly growing, the sooner you can notify engineering, coordinate your agents, and give customers one consistent message.

A scattered problem vs. a spike

Support problems come in two shapes:

  • A scattered problem: one customer complains about it every few days. It should be logged and seen in your weekly review.
  • A spike: within a short window, the number of reports rises above normal. It needs an immediate response.

To spot a spike you need to link conversations to a problem, because you can only count once you know these five conversations are the same problem. That's exactly what you do in issue tracking.

A good alert has three parameters

1. Threshold

How many conversations count as an "incident"? There's no universal number; it depends on your team's volume. A simple way to start: see how many conversations a typical problem gets on an ordinary day and set the threshold a bit above that. After a few weeks, tune it against what you've seen.

2. Time window

A threshold without a window means nothing. "10 conversations" in 15 minutes is a spike; over two weeks it's an ordinary day. A short window suits problems that blow up fast (gateway outages, SMS delivery), while a longer window suits problems that creep upward.

3. Cooldown

After the first alert, if the same problem rings again every few minutes, the team learns to ignore alerts. The cooldown defines how long, after an alert, no new alert should fire for the same problem.

Where the alert shows up

(The issue tracking page covers issues in more detail.)

In Hodhod you can set a threshold, a window, and a cooldown for each issue. When the number of linked conversations in the window crosses the threshold, an in-app alert (a banner or toast) is shown to the owner and admins. A few things worth knowing up front:

  • The alert is not sent by email or SMS and doesn't go to Telegram; it's visible inside the panel only.
  • If you want it to reach another channel (a chat group, an on-call tool, and so on), you can configure a webhook to any https endpoint and route it from there to wherever you like.
  • There's an acknowledge button so the rest of the team knows someone has picked up the problem.
  • The alert ends on its own; you don't need to clear it by hand.

What to do after the alert

A five-step routine, written down before the first incident, helps a lot:

  1. Acknowledge. One person becomes the incident owner and acknowledges the alert so two people don't chase it at once.
  2. Verify it's real. Read a few recent conversations. Are they really one problem, or several problems close together?
  3. Notify engineering. With a clear description and the number of conversations. If you use Jira, connecting support to Jira makes this step easier.
  4. Set one message. All agents give the same answer. Canned responses exist for exactly this.
  5. Announce the resolution. Customers who have been waiting need to know the problem is fixed.

Proactive customer notices

Hodhod has two optional features for this, and both are off by default; you have to turn them on yourself:

  • A widget banner for a problem that's still ongoing, set per inbox. Customers see that you know about the problem before they write, and may not send a message at all.
  • A one-time message to open or pending conversations linked to the issue, sent when the problem is resolved. If you close the issue manually, you confirm the recipients, and no message goes out without your review.

Weigh these against your type of business. For a big outage that everyone can see, a banner usually pays off; for a small problem affecting a limited group, you may be better off giving individual replies.

Common mistakes

  • A threshold that's too low. The alert fires every few hours and the team stops looking.
  • A threshold that's too high. The alert arrives when the incident is already at its peak.
  • An overly broad issue. If "payment problem" is one issue, the alert doesn't tell you which payment problem.
  • No owner. An alert with nobody to acknowledge it is just a colored bar.
  • No review. After each incident, compare your thresholds with what you saw.

Frequently asked questions

Does the alert reach email or Telegram?

No. The alert is in-app and shown to the owner and admins. If you want another channel, you can configure a webhook to any https endpoint.

Do we need to set a threshold for each issue separately?

Yes. The threshold, window, and cooldown are per issue, because different problems have different normal volumes.

Is the customer notice automatic?

No. It's off by default and you have to turn it on. When you close an issue manually, you also confirm the recipients.

Does this replace technical monitoring?

No. This alert comes from customer conversations, so it sees the problems customers have felt. Technical monitoring is still necessary, and the two complement each other. For how this relates to response-time commitments, also read SLA in customer support.