# pdatum API, v1

Job postings and the employers behind them, as an API. This is the data
jobwolverine.com and rxraven.com are built on.

Base URL: `https://pdatum.pearachute.com/api/v1`. This page is served at
`/api/v1/docs` and needs no key.

## Authentication

Every endpoint except this page needs a key:

```
Authorization: Bearer pdatum_...
```

Keys are issued by hand for now. A key is shown once, when it is issued; we
store only a hash of it, so a lost key cannot be recovered, only replaced.
Keep keys out of browsers: responses carry no CORS headers, so a web page on
another site cannot use one.

A key carries scopes. `jobs` reads `/jobs*`, and `employers` reads
`/employers*`. `GET /me` shows which scopes yours has.

## Responses

Every response is JSON. A list looks like this:

```json
{"data": [...], "next_cursor": "WzE3MjcwMDAwMDAsIDQyXQ", "total": 8608}
```

- `total` is how many records matched the filters, across every page.
- `next_cursor` is `null` on the last page. To get the next page, pass it back
  as `cursor=` with the same filters. Cursors are opaque, so don't build or
  edit one.

A single record is `{"data": {...}}`.

An error looks like this:

```json
{"error": {"code": "invalid_parameter", "message": "posted_since: expected a whole number, got 'yesterday'"}}
```

| Status | `code` | Meaning |
|---|---|---|
| 400 | `invalid_parameter`, `invalid_cursor` | A parameter could not be read. Unknown parameters are refused, not ignored. |
| 401 | `missing_key`, `invalid_key`, `revoked_key`, `expired_key` | The key is absent, wrong, revoked or expired. |
| 403 | `insufficient_scope` | The key lacks the scope this endpoint needs. |
| 404 | `not_found` | No such endpoint or record. |
| 429 | `rate_limited` | Slow down; honour `Retry-After`. The limit is 5 requests a second per key, in bursts of up to 60. |

All times are epoch seconds, UTC.

## Jobs

### `GET /jobs`

Open jobs, newest first.

| Parameter | |
|---|---|
| `brand` | `jobwolverine` or `rxraven`. Omit it for both. |
| `employer` | An employer slug, from `/employers`. |
| `q` | Text in the position or the company name. |
| `location` | Text in the location. |
| `remote` | `true` or `false`. |
| `posted_since`, `posted_before` | Epoch seconds. These filter on the posting date, which is the first-seen date when the source gave none. |
| `detail` | `summary` (the default) gives a description excerpt. `full` gives the whole text, capped at 100 records a page. |
| `limit` | 1 to 500 (100 with `detail=full`). The default is 100. |
| `cursor` | From the previous page. |

A job:

```json
{
  "id": 48213,
  "position": "Senior Data Engineer",
  "company": "Acme Corp",
  "employer": "acme",
  "location": "Boston, MA",
  "remote": false,
  "salary_min": 150000, "salary_max": 180000, "salary_currency": "USD",
  "tags": ["python", "senior"],
  "description": "...", "description_is_excerpt": true,
  "application_url": "https://...", "source_url": "https://...",
  "brand": "jobwolverine",
  "open": true,
  "posted_at": 1727000000,
  "first_seen_at": 1727003600,
  "last_seen_at": 1727500000,
  "updated_at": null
}
```

Field notes:

- `employer` is the slug of the employer record, or `null` when the company
  name has no record of its own.
- `posted_at` is `null` when the source gave no posting date. `first_seen_at`
  is always when we found the job.
- `last_seen_at` is the last time a crawl still listed the job.
- Salaries are whole units of `salary_currency`. Most jobs don't state a
  salary, and both fields are `null` then.

### `GET /jobs/count`

Takes the same filters as `/jobs` and returns `{"data": {"count": n}}`. It's
cheap, so use it to size a query before pulling the results.

### `GET /jobs/{id}`

One job, open or closed, with its full description.

### `GET /jobs/changes?since=T`

Every job added, closed or reopened after `T`, oldest change first. Closed
jobs are included with `"open": false`, and each record carries `changed_at`.
It takes `brand`, `detail`, `limit` and `cursor`.

This is how to keep a copy in sync without pulling everything again:

1. Pull `/jobs` once, and note the time you started.
2. From then on, call `/jobs/changes?since=<last changed_at you saw, minus 1>`.
   Follow `next_cursor` to the end, and upsert each record by `id`.

Times are whole seconds, so the "minus 1" means a record may arrive twice. It
never means one is missed.

## Employers

### `GET /employers`

Employer records, in id order.

| Parameter | |
|---|---|
| `q` | Text in the name. |
| `domain` | An exact domain, e.g. `pfizer.com`. |
| `brand` | Only employers with open jobs on that brand. |
| `limit`, `cursor` | As for `/jobs`. |

An employer:

```json
{
  "id": 17, "slug": "acme", "name": "Acme Corp", "domain": "acme.com",
  "origin": "postings", "created_at": 1726000000,
  "open_postings": {"jobwolverine": 12, "rxraven": 0},
  "hiring": true
}
```

`hiring` is `true` when there are open jobs, or when our own check of the
employer's hiring pages found some. It is `false` when that check found hiring
pages with nothing open. It is `null` when we can't tell, which is not the
same as not hiring.

### `GET /employers/{slug}`

One employer. On top of the list fields it carries:

- `aliases`: the company names that resolve to this employer.
- `facts`: every fact we hold, each as
  `{key, value, source, method, observed_at}`.
- `history`: every value those facts have had, oldest first, in the same
  shape.

Every fact names its source. `method` is which version of our own check
produced the value when the source is ours (`probe`), and `null` for a source
that has no versions. When the method changes, history gets a row even if the
value did not: that row is the new version's starting point, not a change at
the employer. Compare values only within one source and method -- a better
check finding a hiring page is not the employer starting to hire.

## `GET /me`

The key you are calling with: its name, holder, scopes and expiry, plus
`usage` for today and the last 30 days. Usage counts requests, and records
returned.
