Integration•8 min read

How to Monitor Celigo Integrations Before Small Errors Become Bigger Problems

Watch the Business Result, Not Just the Error Count

A few failed orders may look harmless until the warehouse reaches its shipping cutoff. A delayed inventory update can leave a sales team working with old stock levels. Good monitoring gives someone time to act before either issue affects customers.

We conducted a documentation review of Celigo, Google SRE, Oracle, and AWS guidance to build this monitoring routine. Our analysis focuses on three questions: did the work run, did the right records arrive, and can the team recover safely? The thresholds below are examples to adapt, not findings from a live customer test.

Start With a Clear View of Each Important Flow

Celigo says its integration dashboard shows flow runs, status, timing, and record counts, including errors. It also provides tools to review and retry errors. Available history depends on retention limits. [1] Use that view as the starting point for daily checks, then confirm key results in the destination system.

For every business-critical flow, record its owner, backup, schedule, expected volume, and latest acceptable delivery time. Include dependencies. An order flow may depend on customers or items arriving first. Keep this short enough that someone covering a shift can use it.

  • Orders: Who checks that paid orders reach NetSuite before the shipping cutoff?
  • Inventory: How old can a stock update be before a channel owner needs to act?
  • Payments: Who checks missing settlements before finance closes the day?
  • Master data: Which downstream flows depend on customer, item, or price updates?

Track Early Warning Signs Alongside Failures

Google's SRE book names latency, traffic, errors, and saturation as four core monitoring signals. [2] For a Celigo process, we suggest translating those into delivery time, record volume, failed records, and pressure on processing capacity. You may need source reports, destination searches, or an external monitor to build this view; these are not all guaranteed fields in one Celigo dashboard.

Suggested signals for an integration monitoring routine
SignalWhat to compareWhen to investigate
Last expected runActual start time against the agreed scheduleA run is missing after its allowed delay
Delivery delaySource event time against destination arrival timeRecords approach the business deadline
Oldest unresolved errorError age against its owner's response targetOne record stays blocked while newer work succeeds
Record volumeEligible source records against confirmed destination recordsAn unexplained gap remains after normal processing time
Run durationCurrent duration against similar days and volumesRuns repeatedly take longer or overlap
Repeat failuresError type, affected step, and record key over timeThe same issue returns after a retry or repair

Compare like with like. A quiet Sunday is a poor baseline for a promotion day. Start with a normal trading period, note peak periods separately, and revisit the baseline after schedule or mapping changes. Keep the original timestamps and use one time zone when comparing systems.

Set Up Notifications and Test the Handoff

Celigo's instructions say to open an integration's Notifications tab, choose all or selected flows under flow-error notifications, select the connections to watch for offline events, and save. [3] Check subscriptions for both the primary owner and the backup. A shared expectation that “someone gets the email” is not a handoff.

Celigo documents a 15-minute check for new open and newly resolved errors before sending the relevant email summary. It also notes that some errors resolved before a notification is triggered do not generate an email. [4] Treat email as one signal, rather than proof that every scheduled run happened or every record arrived.

Test delivery through a controlled failure in an isolated test setup. Confirm who received the message, whether they could find the affected flow, and who takes over if they are unavailable. If a business deadline is shorter than the documented notification interval, add an independent freshness check with a suitable polling interval.

Use Business Deadlines to Set Alert Levels

A useful alert tells someone what is at risk and what to do next. The example below is a suggested policy for an order flow expected to run every 15 minutes. Adjust it for normal runtime, operating hours, and the actual shipping deadline.

  • Review: A single invalid address blocks an order. Send it to the order owner with the record key and correction steps.
  • Investigate: No expected run has started 10 minutes after its scheduled time. Have the integration owner check the schedule and connections.
  • Escalate: Eligible orders remain undelivered as the shipping cutoff approaches. Notify the operations lead and agree on a recovery plan.

Include the environment, flow name, affected step, oldest record age, business impact, and owner in the incident note. Link to a short recovery guide. Avoid copying full customer records or credentials into notification messages.

Follow a Repeatable Response Path

When an alert arrives, first check its scope. One bad record needs a different response from an offline connection affecting several flows. Use this five-step path to keep diagnosis, recovery, and business confirmation together.

DetectFind the missed run, delay, or error.
AssessCheck affected records and deadlines.
AssignName an owner and next update time.
RecoverFix the cause and retry safely.
ConfirmVerify delivery and watch the next run.

Keep the incident open until the intended business result is confirmed. A cleared error queue alone is not enough. Check the destination record, related downstream steps, and any work that built up during the failure.

Check NetSuite Capacity When Runs Slow Down

Oracle says NetSuite's Concurrency Monitor tracks web services and RESTlet performance, with overview and detailed dashboards. [5] When those calls slow down or fail, compare the time of the Celigo issue with NetSuite's concurrency data. This helps you investigate whether other integrations are competing for capacity.

Do not increase parallel requests just because a backlog is growing. First check whether the delay comes from the source export, a lookup, the destination import, or a dependent process. Record the step and time window so the next person can repeat the investigation.

Make Retries Part of Recovery, Not the Whole Plan

AWS explains that idempotent API design makes repeated requests safer by avoiding repeated side effects. [6] That principle matters when a timeout leaves you unsure whether an order or payment was created. It does not prove that your Celigo flow is safe to replay.

Before retrying, check the destination for the record's stable business key. Review whether the flow creates, updates, or upserts records and how it matches existing data. Fix invalid input or access problems first. Where practical, validate a small recovery batch before replaying a larger queue.

For example, imagine 12 orders fail because an item code is missing. Adding the item may remove the immediate cause. Still, confirm whether any orders were partially processed and whether later steps ran. Then verify the recovered orders in the destination. This is an illustrative scenario, not a reported customer result.

Build a Routine the Team Can Sustain

During each staffed operating period, review critical flows, missed runs, aging errors, and delivery gaps. For overnight processes with tight deadlines, arrange coverage and automated escalation that match the business need. A morning review cannot protect a midnight cutoff.

Each week, review recurring error types and alerts that needed no action. Fix the most disruptive recurring cause and adjust noisy thresholds only after checking what they would hide. After any flow change, confirm the schedule, subscriptions, owner, and destination checks still fit.

Keep a short incident history outside any limited run-history window. Save the cause, affected record keys, recovery steps, and confirmation time under your team's access and retention rules. For a broader routine, see our Celigo integration maintenance checklist. If the same failures keep returning, use our Celigo integration audit guide to review the design.

References

  1. Celigo Help Center: Explore the integration dashboard.
  2. Google SRE: Monitoring Distributed Systems.
  3. Celigo Help Center: Get error notifications via email.
  4. Celigo Help Center: How integrator.io determines when to email an error notification.
  5. Oracle NetSuite: Concurrency Monitor Overview.
  6. Amazon Builders’ Library: Making retries safe with idempotent APIs.

Build a Celigo Monitoring Plan Your Team Can Use

We can review your critical flows, identify gaps in alerts, and define recovery steps around your business deadlines.

Frequently Asked Questions

Practical answers about monitoring Celigo runs, alerts, and recovery.

What should I monitor in Celigo each day?

Check critical flow runs, connection status, unresolved error age, run duration, and whether expected records reached the destination. Start with flows tied to shipping, stock, or finance deadlines.

Does Celigo send error notifications immediately?

Celigo documents a 15-minute check for new open and newly resolved errors before sending relevant email summaries. Some errors resolved before a notification is triggered do not produce an email. Use a separate freshness check when your deadline requires faster detection.

How do I turn on Celigo email alerts?

Open the integration’s Notifications tab. Select the flows for flow-error notifications and the connections for offline notifications, then save. Confirm subscriptions and test that the owner and backup receive the expected messages.

Can a flow look successful while records are missing?

A successful run does not by itself confirm that every expected business record was selected and delivered. Compare eligible source records with destination records, accounting for filters, timing, and exclusions.

How should I choose alert thresholds?

Start with the business deadline, normal run schedule, and usual processing time. Set a warning early enough for an owner to act. Treat example thresholds as starting points and adjust them using your own operating history.

Should I retry every failed record automatically?

No. Check the cause and whether the destination already contains the record. Fix data or access problems first, and confirm that replaying the operation will not create duplicate records or repeat a business action.

What helps diagnose slow Celigo flows connected to NetSuite?

Find the slow flow step and compare its time window with NetSuite’s Concurrency Monitor for web services and RESTlet activity. Review competing work before changing schedules or increasing parallel requests.

When is an integration incident resolved?

Close it after the affected records reach the correct destination, downstream work is checked, and the next expected run is healthy. Record the cause, recovery steps, and owner of any follow-up fix.