use cases

Matching Candidates to Jobs and Deleting Them Cleanly: A Small Recruiting Agency on HyperCrux

Say a four-person agency places data engineers and analysts, and holds about 8,000 CVs. When a client sends a new job, the first move is a shortlist: the people whose experience is closest to the job description, among those who are available and haven’t already applied. Pitching someone for a job they applied to last week is a quick way to look careless in front of a client.

The agency also holds personal data, and that comes with duties. Under data protection law a candidate can ask to be deleted, and the agency’s own policy says CVs go two years after the last contact. Either way the CV has to go along with everything made from it, the vector included, and every link that ties the person to jobs and clients.

A job’s closest candidates by distance, with the one who already applied left out. Beside them, deleting a candidate removes the record and its links at once, and a retention rule deletes candidates nobody has contacted in two years.

The Shortlist

Each candidate is a record with their details, an available flag, the date of last contact and a vector made from the CV. Each job has a vector made from its description. When a candidate applies, an applied link goes from them to the job. Then the shortlist is one call:

hits, err := db.Nearest("candidate", jobVec, 20,
	"available = 1 AND key NOT IN (SELECT value FROM json_each(walk(?, 1, 'applied', 'in')))", jobKey)

jobVec is the job’s stored vector. The filter is plain SQL, and it can use walk like any other query: here it walks backwards from the job along applied links and leaves out everyone it finds. The twenty closest of the rest come back in order.

A Deletion Request

err := db.Delete("candidate:4412")

That one call deletes the record and its vector, which lives in the same row, and the triggers in the file delete every link to or from it in the same transaction. There’s no second system to purge, and no moment when the CV is gone but a search can still find it.

The Two-Year Rule

The retention policy is one SQL statement, run once a month:

res, err := db.Exec(`DELETE FROM candidate WHERE last_contact < date('now', '-2 years')`)

It’s plain SQL, and the triggers still run for every row it deletes, so each candidate’s links go with them. res.RowsAffected() gives the count for the agency’s records, and db.Check() afterwards confirms that keys, rows, links and vectors still agree.

SQLite leaves deleted bytes in free space inside the file until it needs that space again. After a deletion request or the monthly purge, rebuild the file and fold the write-ahead log back in:

_, err := db.Exec("VACUUM")
if err == nil {
	_, err = db.Exec("PRAGMA wal_checkpoint(TRUNCATE)")
}

VACUUM writes the rebuilt pages to the log first, so the checkpoint is what finally replaces the old pages in the file. Backups taken earlier still hold the old data, so they need a retention rule of their own.

Speed

In the recorded benchmarks on a two-core cloud machine, a search among 10,000 vectors of 384 values took about 42 milliseconds. The agency’s 8,000 CVs are fewer than that, and a delete is a single small transaction. The file lives on the server that runs the agency’s own tools, and any number of processes there can share it.

People Still Decide

The ranked list is where a recruiter starts reading. Embeddings pick up whatever the CVs and job ads happen to emphasise, so a person still decides who gets the call. In some places, the EU among them, software that ranks job candidates comes with legal duties of its own, and an agency using it should know what those are.

Matching and forgetting pull in opposite directions. One wants data kept and connected, the other wants it gone without a trace. With the vector in the candidate’s own row and the links tied to it by triggers, both come down to a single call. The reference lists Nearest, Delete and the rest of the Go API.