Skip to content

Log Investigation

Most access tickets arrive with a time attached: “it did not work this morning at about 10:40”. Answering means opening the web logs, then the firewall logs, then ZPA, then checking whether anyone else had the same problem, and lining up the results by hand. The Log Investigation does those reads for you, one after another, and returns one timeline with diagnosis cards that name the likely cause and the next step.

It is for the help desk engineer holding a ticket, and for anyone who has to explain afterwards what happened at a given moment.

When to use it

  • A ticket with a time: a user could not reach a host or a cloud app at a known moment.
  • A ticket with only a day: “sometime yesterday”. Only the day finds the moment for you.
  • A failed request in a HAR: jump from the request straight to an investigation at that exact time.
  • “Is it just me?”: the investigation checks whether other users failed on the same target in the same window.

Where to find it in the UI

Open the Troubleshooting Engine and go to the Investigation tab. The user and the destination come from the scenario selected in the rail, so build the criteria first.

You can also start from the HAR Analysis panel: Investigate this request on a failed request opens the tab with the host and the time of that request already filled in.

Step by step

  1. Select the scenario in the rail. The user is taken from it and is not editable here.

  2. Choose the target. When the scenario has a cloud app, choose between the cloud app as a whole and a host. For a host, type it in the host field: a plain host name is searched as an exact match, and a leading *. includes every subdomain.

  3. Choose when. Three ways, depending on what the ticket says:

    OptionUse it when
    A few minutes ago (5, 15 or 60 minutes)The user is on the phone right now
    At a given time, with the timezoneThe ticket has a time
    Only the dayThe ticket has a day and no time (see below)

    Set the margin before and after the moment: 1, 5, 15 or 30 minutes, 5 by default. A wider margin catches more, and costs more time on each read.

  4. Choose the steps. Each step is one log read, run in this order:

    StepWhat it reads
    User to hostThe user’s traffic to the target. Always on
    ZPA sessionsThe user’s ZPA sessions towards the host
    Other failures of the userEverything else that failed for the same user in the window
    Sub-resourcesThe other hosts the page loads, where the real failure often hides
    Firewall sessionsThe user’s firewall sessions in the window
    Other usersWhether other users failed on the same target

    ZPA sessions and Sub-resources need a host, so they are not offered when the target is a cloud app.

  5. Press Run investigation. Nothing runs before you do. The reads go one after another and the timeline fills in as they complete.

  6. Read the diagnosis cards, then the timeline behind them.

The Investigation tab before a run: the user from the scenario, the destination set to the cloud app OneDrive instead of a host, When set to A few minutes ago (15 minutes) with a 5 minute margin, the steps with ZPA sessions and Sub-resources unavailable because they need a host, and the Run investigation button

Reading the result

The result is one timeline that merges every read, in time order, and a set of diagnosis cards on top. Each card has a title that names the cause, a summary of the evidence, and a next step. Show these events in the timeline jumps to the events behind the card.

The causes the cards recognise:

AreaDiagnosis
CertificatesThe application does not trust the inspection certificate; the destination presented a certificate Zscaler does not accept
PolicyBlocked by the SSL/TLS inspection policy, by URL filtering, by cloud application control, by data loss prevention, by file type control, by the firewall
Threat protectionBlocked or held by the sandbox; blocked as a threat; dropped by intrusion prevention
PerformanceThe destination server answered with an error; the destination server was slow; Zscaler processing was slow
ZPAZPA sessions failed

When a failure does not fit any of these, the card says so (“Failed for a reason ZHERO does not classify yet”) and the timeline still shows the raw events. The Other users step adds how many other users had failures on the same host in the window.

The result of an investigation: diagnosis cards (an application that does not trust the inspection certificate, a block by cloud application control) each with hosts, time range, next step and Show these events in the timeline, the note that only one user had failures on the host, and the single timeline of web and firewall events with outcome, marks, reason and rule

Only the day

When the ticket gives a day but no time, choose Only the day, pick the day of the ticket in the calendar and its timezone, then Load the day. ZHERO reads the user’s traffic towards the target for that day in one pass and draws it as a heatmap of 144 cells, ten minutes each. Cells that contain a block are marked in red.

Above the heatmap, Where to look first suggests the moments that usually matter: the first block of the day, the busiest half hour, the moment traffic stops. Click a cell to select a 30-minute window around it (Widen to 1 hour if you need more), then Investigate this window to run the investigation there.

Only the day needs a scenario with a ZIA user. If the day holds more traffic than a single read returns, the page says so and offers Read the rest.

Only the day for a host: the day and its timezone, Load the day, a heatmap of the whole day with the blocked cells marked in red, the Where to look first shortcuts (first block, busiest half hour, when traffic stops), a selected 30 minute window with its counts, and Investigate this window

Saving and naming

Every investigation is saved in the History sub-tab, next to Investigation. From there:

  • Restore reopens an investigation with its results, without running any query.
  • The pencil gives it a name of your own (up to 80 characters).
  • The pin keeps it: unpinned entries are kept for 7 days, up to the latest 10.
  • An investigation that did not finish is saved as Interrupted and can be picked up with Resume.

The History sub-tab: saved investigations with their window, rows and read time, one renamed "Claude API blocked, ticket 4821" and pinned, each with Restore, pin and delete

What to watch

  • The cards before the timeline. The timeline is the evidence; the cards are the reading. Start from the card, then check its events.
  • Sub-resources. “The site does not work” often means the page loaded and one of the hosts it depends on did not. That is the step that finds it.
  • Other users. One user failing is a user problem; twenty users failing on the same host is a policy or destination problem.
  • Certificate cards. An application that does not trust the inspection certificate is fixed with an SSL inspection exemption or by deploying the certificate, not by changing URL filtering.
  • The margin. Start narrow. Widen only if the timeline is empty around the moment.

Limits and notes

  • The Investigation is part of the Troubleshooting Engine and needs the full licence.
  • One log query at a time. Zscaler allows one log job per admin session, which is why nothing runs until you press Run: an automatic read would replace the one you already have running.
  • Each step reads up to 5,000 rows. On a very busy user, narrow the margin to keep the window complete.
  • A read that stops advancing is stopped after a few minutes, with a message, rather than left hanging.
  • The logs are Zscaler’s. The investigation reads them through your authenticated session, as the console would; what your admin role cannot read, it cannot read either.
  • History is stored locally in your browser.

FAQ

Why do I have to press Run? The form is already filled in. Because Zscaler runs one log query per admin session. Starting one automatically would silently cancel whatever you were already running, in the drawer or in a Logs panel.

Can I investigate a whole cloud app instead of one host? Yes, when the scenario’s destination is a cloud app. The steps that need a single host (ZPA sessions and sub-resources) are then not available.

The timeline is empty. Check the timezone and the margin first: an investigation at the right time in the wrong timezone reads an hour where nothing happened. If the ticket is vague, use Only the day.

Does the investigation change anything in the tenant? No. It only reads logs.

Next steps

  1. Take a ticket with a time and run the investigation with the default margin
  2. Read the first diagnosis card and follow it into the timeline
  3. Next time a ticket says only “yesterday”, try Only the day
  4. Name the investigation you want to keep, and pin it