Skip to main content

Opening service-desk tickets from a Console feature

backend/internal/ticketing opens and follows up SUP tickets for any Console feature. Maintenance windows are its first caller. Use it instead of calling JSM from a handler.

It keeps its own outbox: the tables ticket_outbox (one row per source's ticket) and ticket_outbox_ops (that ticket's ordered create / comment / update-fields ops), added by migration 000260_ticketing_and_maintenance_notify. Do not confuse them with service_desk_tickets, the pre-existing read model of JSM tickets that data sync maintains: the outbox holds what Console intends to send to JSM, the read model holds what JSM already has.

The rules​

Summary and Description are customer-visible

A ticket is raised on the client's service desk organization, so its Summary and Description show on the client's portal. Put nothing internal in them: no hosts, targets, alerts, monitors, creator, internal ids or Console URLs. Build them from the same narrow, client-safe view you would send to a client chat (maintenance uses ClientWindow).

Internal detail — a Console link, who did what, what is affected — goes in a comment: Comment always posts an internal comment (public=false). Maintenance enqueues its Console link as an internal comment right after Open, in the same transaction, with its own dedup key.

  1. Enqueue in your own transaction. Service.Open, Comment and UpdateFields take the caller's sqlx.ExtContext. The ticket intent commits with the row that caused it, or not at all. Saving never calls JSM inline. The worker does that, every 10 s, on every replica. Your store needs a seam that runs callbacks inside its write transaction; copy store.MaintenanceTxHook (the maintenance store runs its hooks after the row is written and before commit, on every write path, so a hook sees the new id and updated_at).
  2. One ticket per source. A source is (kind, id), for example ("maintenance_window", <window id>). Open is idempotent per source; a source whose ticket belongs to another client is refused with an error wrapping domain.ErrConflict. Comment and UpdateFields are no-ops for a source without a ticket. Give every follow-up a dedup key that names the change (window:{id}:rev:{updated_at}:comment), so a retried request cannot enqueue it twice.
  3. Cap comment bodies yourself. Pass ticketing.Truncate(body, ticketing.MaxCommentLen) to Comment. An empty or over-length body is an error, checked before the ticket is looked up, and it is returned inside your transaction, so it rolls back the save that caused it (even for a source with no ticket at all).
  4. Fields are typed. Set ticketing.Fields{StartDate, DueDate, MaintenanceWindow, Labels}. Never write a customfield_* id: they live in one table in ticketing.go. The organization (customfield_10002) always comes from the client's org mapping, never from you. Dates are the local date of the time.Time you pass, so pass them in the zone they should be read in. UpdateFields writes both dates every time: a nil date clears it, so UpdateFields(…, ticketing.Fields{}, …) clears the start and due dates. Send the full set. Fields is a closed set: a new Jira field is added to Fields and its id to the table in ticketing.go, never passed in from a feature.
  5. Register your source in the router (ticketRegistry.Register(ticketing.SourceSpec{…})). Register refuses a spec without a Kind, a non-empty ReadPerm or an Authorize, and the router panics on a registration error, so a wiring mistake fails at startup. Registering also adds the kind to the source_kind label set of proxima_ticketing_ops_total; there is no hand-kept list to update. Set WritePerm to the permission that edits your rows (maintenance:write for maintenance): GET /clients/{id}/ticketing/availability?source_kind=<kind> then accepts it on the client instead of servicedesk:read, so your form can offer the option to the people who edit the row. Without source_kind the endpoint needs servicedesk:read. The route's middleware takes its permission set from the registry, so nothing else changes.
  6. Bind every read to the authorized client. Service.Get(ctx, clientID, src) and Service.Retry(ctx, clientID, src) take the client your own authorization resolved for the source, never one taken from the request. A ticket of another client answers domain.ErrNotFound, exactly like a missing one.
  7. Embed the ticket in your own DTO (ticket: TicketDTO | null via Service.Get). There is no ticket list: your read permission is the ticket's read permission.
The Authorize contract: invisible means not found

SourceSpec.Authorize(ctx, ac, id) answers "which client owns this row, and may this caller see it". The generic retry endpoint (POST /api/v1/ticketing/{kind}/{id}/retry) turns its answer into the response, so it decides whether a row's existence leaks.

  • A row the caller may not see — missing, deleted, another client's, or outside the caller's environments — must return (_, false, nil) or an error wrapping domain.ErrNotFound. Both become the same 404 as an unknown kind.
  • Return any other error only for a genuine failure (the database is down). It is answered with its own status, so returning domain.ErrForbidden, ErrConflict or similar for an invisible row tells the caller that the row exists.
  • ok = true must mean the caller may read the row. The endpoint then also checks ReadPerm and servicedesk:write on the returned client.
  • An Authorize that hides deleted (or archived) rows makes their tickets un-retryable: Retry answers 404, and a failed op on such a ticket stays failed. That is the maintenance window's behaviour; decide it deliberately for your kind.

The maintenance window's AuthorizeTicketSource is the reference implementation.

What the worker does​

  • create:

    1. Resolve the client's org mapping. If there are several, it takes the first by organization name.
    2. Search for a ticket already labelled console-src-<source id> and adopt it. A create that JSM accepted but Console never recorded (a crash between the two) is adopted instead of duplicated.
    3. Otherwise record on the op, committed, that a create is being attempted (create_attempted_at), then CreateRequest with the request type PROXIMA_TICKETING_REQUEST_TYPE_ID, raised on behalf of requested_by (the integration account if that user cannot be resolved). If the request type's form rejects labels, it creates without them and the edit adds them. Detecting that rejection relies on JSM naming labels in its 400 message; this is unverified against production JSM, and a differently worded rejection fails the op permanently rather than creating anything.
    4. Record the key.

    Never a second ticket. If a create was already attempted, no key was recorded and no labelled ticket is found, the op fails at once with "Jira: a previous attempt may already have created this ticket — if it is in SUP, add the label console-src-<source id> to it, then press Retry; if not, just press Retry." (with the real label). It never creates again: an unlabelled issue from a lost response is invisible to the label search. A person looks in SUP. If the issue is there, they add that label to it before pressing Retry, so Retry adopts it; a bare Retry would create a second ticket. After adding the label in SUP, wait about a minute before pressing Retry, so Jira's search can see it. If it is not there, they just press Retry. Retry clears the marker. A definite JSM rejection (any 4xx, 429 included) clears the marker at once, because nothing was created. 5. Set the fields the portal form does not carry with a Jira issue edit. The op stays pending until this edit succeeds; a re-claimed create op that already has a key only redoes the edit.

  • comment: an internal comment (public=false).

  • update_fields: writes both dates (a missing one is cleared) and adds labels.

  • Order: ops on one ticket run strictly in order. A comment waits until the ticket has a key.

  • Claims and leases: each tick claims up to 20 due ops with FOR UPDATE SKIP LOCKED under a 5-minute lease. The lease is also a fencing token: completing, retrying or failing an op matches the claim's lease, so a late outcome from a worker whose lease ran out is a no-op and cannot overwrite the result of the replica that re-claimed it. An op starts only while enough of its lease is left for its worst case (claims too close to expiry are skipped and picked up again after the lease ends), and it runs under a deadline 30 s before the lease ends.

  • Retries: 30 s × 2^n backoff (30 s after the first failure, then doubling), capped at 30 min. After 8 attempts the op fails. A JSM 4xx other than 429 fails at once.

  • Delivery is at-least-once. A comment JSM accepted whose completion was not recorded (a crash, a lost database connection, a shutdown) is posted again when the op is re-claimed. Write comments so a duplicate is harmless.

  • Failures: a failed create fails the ticket, and POST /api/v1/ticketing/{kind}/{id}/retry re-arms it. Readable reasons include "Ticket creation is not configured", "Service desk is not configured", "No service desk organization is mapped to this client" and Jira: <JSM's message>. Console-side failures read as Console: … and never carry internals.

  • A failed ticket can still have an issue key. If JSM created the issue but the field edit after it failed, the ticket is failed and carries issue_key. Show and link the key in that state too; Retry only redoes the edit.

Frontend​

frontend/src/components/ticketing/ holds the reusable pieces:

  • useTicketingAvailability(clientId, sourceKind?)
  • <OpenTicketOption checked onChange clientId sourceKind help keptNote keptNoteUnavailable />: the checkbox, with the reason it is unavailable. Pass your sourceKind so your WritePerm reads availability. keptNote shows under a ticked box the caller cannot tick again for lack of servicedesk:write; keptNoteUnavailable when the client's ticketing is unavailable.
  • <TicketChip ticket sourceKind sourceId clientId onRetried />: the key linking to the portal, "Creating…", or "Ticket failed" with Retry behind servicedesk:write. Pass an onRetried that invalidates your page's query, so the chip shows the re-armed state.

One request type​

Every Console-opened ticket uses the one request type PROXIMA_TICKETING_REQUEST_TYPE_ID (the SUP "Task" type). There is no per-source override; a feature that needs a different request type needs that added to ticketing first.

Rolling out​

Ticket creation stays off until PROXIMA_TICKETING_REQUEST_TYPE_ID is set; until then every create fails with "Ticket creation is not configured". Before setting it, confirm that the request type accepts labels on create: GET /rest/servicedeskapi/servicedesk/{id}/requesttype/{requestTypeId}/field must list labels. If it does not, tickets are still created (without the label, which the edit adds), but a create whose response is lost — a timeout, a crash, a deploy — cannot be found again by label. That op then fails naming its console-src-<source id> label: a person looks in SUP and, if the issue is there, adds that label to it before pressing Retry.

The same manual step applies on any request type during a JSM incident: a create that hit a 5xx or a timeout may or may not have created the issue, so its retry fails the same way unless the label search finds it, and needs a manual Retry.

Metrics​

  • proxima_ticketing_ops_total{source_kind,kind,result}
  • proxima_ticketing_backlog{state}
  • The JSM client's proxima_jsm_request_* metrics still apply. See docs/standards/metrics.md.