PROJECT STORY — THE SECOND BUILD

I DIDN’T WANT
ANOTHER RESUME GENERATOR.

A project story about truthful representation, judgement, controls, real use — and learning where AI should not be used. Written by Sangharsh Redhu, who built Joblinko as a second AI-assisted project and kept a record of what went wrong.

01

WHY THE SECOND PROJECT HAD TO BE DIFFERENT

Verlinko was the first time I took an idea all the way into a substantial working system using AI. It taught me a lot, but not always in the way I expected.

I learned how quickly AI could create something that looked finished. I also learned how easily I could mistake that for something tested, reliable, or even correctly understood.

I accumulated too much context. I allowed old and new decisions to coexist. I trusted reassuring names before checking behaviour. I built things that measurements later told me to remove. By the end I was working very differently from the way I started.

So for the next project I wanted to test something else. What happens if I apply those lessons much earlier? Can I define the source of truth first? Can I separate what needs judgement from what ordinary software does more reliably? Can I build rules into the workflow instead of leaving them inside documents? Can I let AI help without giving it permission to quietly change the facts?

02

THE PROBLEM WAS NEVER WRITING A RESUME

I did not start Joblinko because the world needed another AI resume writer. By then it was obvious that generative AI could produce a polished resume, write a cover letter, and summarise a job description. That was almost the least interesting part of the problem.

A real job search is messier. You begin with a career spanning ten, twenty, thirty years. The person remembers some achievements clearly and others badly. Important numbers may never have been written down. Old resumes contradict one another. The same experience is relevant to several different kinds of role.

Then a job description arrives, and it tells the AI exactly what the employer wants to hear. That creates an obvious temptation: take what the person has actually done and make it sound closer to what the employer wants.

Sometimes that is legitimate tailoring. Sometimes it crosses a line. Someone may have led a transformation without holding the title the employer uses. They may have run a similar process in another industry. They may also simply not have done it.

A model does not feel the discomfort a person feels when saying something technically impressive but not quite true. Its job is to produce a useful answer. That made the real question far more interesting than resume formatting: can AI help somebody represent their career more effectively without quietly creating a more convenient version of that career?

03

I DID NOT START FROM A BLANK PAGE

Joblinko did not begin as an empty repository. I started from an existing open-source project built by another developer for his own job search. It already contained working ideas about job discovery, evaluation, resume generation and testing. It gave me a scaffold instead of forcing me to recreate every basic mechanism.

Using somebody else’s foundation taught me something quickly: you inherit more than code. You inherit assumptions.

And inherited assumptions are harder to see than inherited code, because they rarely announce themselves as assumptions.

The original system had been built for a different person, with a different career and different decisions about how a job search should work. Some fallback role categories buried inside it were still the original author’s AI-engineering categories long after the rest of the project had been adapted to completely different people.

Nothing crashed. That is what made it dangerous. A hidden default can survive for months precisely because it works — it simply works for the wrong person.

Over time I replaced the instruction files, reworked the profile model, changed the evaluation logic, added job sources, built new interview preparation, introduced bounded chat experiences, added stronger validation and changed the release process. The lesson was not that reuse was wrong. Reuse let me move faster. But every inherited assumption had to earn the right to stay.

That feels increasingly relevant to enterprise AI. Organisations will buy platforms, agent templates and prebuilt workflows rather than build everything. The danger is rarely the code they inherit. It is the assumptions they never realised came with it.

04

ONE VERSION OF THE PERSON’S CAREER

One lesson from the first project was that competing sources of truth get expensive fast. So Joblinko has one canonical career history — a human-readable file holding the person’s underlying experience. Everything generated is supposed to trace back to it.

Simple in theory. Not in practice. Most career histories are incomplete, because a traditional resume is built to be concise. It leaves out projects, numbers, problems solved, team sizes, technologies and context that turn out to matter years later.

So the system asks questions instead of asking someone to rewrite their career. What did you actually own? How many people were involved? What changed? What was the starting point? Can that number be supported?

The most useful discovery was that better output sometimes required no improvement to the AI at all. It required better facts. In one enrichment session, eleven genuine facts were added to the underlying history — closing roughly twenty-five gaps that later mattered to matching. No model could have inferred them responsibly, because they were not in the source.

A model cannot be grounded in information that was never captured.

The same is true in operations. If the real exception, the judgement call or the workaround exists only in somebody’s head, an AI cannot reliably act on it either — and no amount of model capability closes that gap.

THE PART THAT CHANGED MY MIND

I STOPPED GUESSING
AND STARTED MEASURING.

The most valuable day of this project produced no new feature at all. It produced numbers — and several of them said the opposite of what I expected.

MEASURED, NOT ASSUMED
REAL RUNS · REAL DOCUMENTSCONCLUSIONS KEPT
BETTER FACTS BEAT A BETTER MODELEleven genuine facts added to one career history in a single sitting closed roughly twenty-five gaps that later mattered to matching. No model could have inferred them, because they were not in the source.THE BIGGEST LEVER
THE CONSTRAINTS WERE DOING REAL WORKAsked to write the best possible resume with the format rules removed and the truth rule kept, the output drifted twice in about 1,200 words — inventing a detail the source never claimed, and borrowing a word straight from the job description. The constrained system drifted zero times on the same inputs.2 DRIFTS vs 0
STRUCTURAL TINKERING DID NOT MOVE QUALITYTwo rounds of rewriting the generation instructions produced documents almost identical to their predecessors once cosmetic differences were discounted. Worth doing for correctness. Invisible to a recruiter.NEARLY IDENTICAL
A CHEAP FILTER, MEASURED BEFORE IT SHIPPEDAcross two real boards and 600 listings, the country filter removed 355 with no false positives found among the drops. That is not proof it will never err — it is why the rule was kept deliberately timid.355 / 600
THE EXPENSIVE CONCLUSIONS WERE WRITTEN DOWN SO THEY WOULD NOT BE PAID FOR TWICE
05

THEN I FOUND THE SOURCE OF TRUTH COULD ALSO BE WRONG

Calling something the source of truth does not make it true.

During one enrichment conversation, an existing entry described the person as having “led the practice.” The conversation established that this was not accurate — the role had been an individual contributor position. The wording was corrected.

That mattered, because an earlier generated resume had already used the stronger version as evidence of leadership. The AI had not invented the claim. It had faithfully used the source it was given.

Grounding removes one class of problem. It does not remove the need to question the ground. An enterprise system grounded in an incorrect policy, a stale customer record or a wrongly coded account status can be perfectly grounded and perfectly wrong. The source of truth needs governance too.

06

TELLING THE AI TO BE TRUTHFUL WAS NOT ENOUGH

There are many instructions in Joblinko telling the model not to invent experience. I do not consider those controls. They are instructions.

The stronger control runs after generation. A deterministic validator examines a generated resume before any PDF exists, looking for unsupported metrics, suspicious credentials, inflated language, structural inconsistencies and other claims checkable against the canonical history. It can warn. It can also block generation completely.

That distinction matters. If something should not be allowed through, “please don’t do this” is a much weaker design than a gate that can actually stop it.

The validator also taught me not to overstate what a control proves. It has known gaps, and one is important: if the career history contains a legitimate metric, the validator can confirm that number exists somewhere in the source — but it cannot yet prove the number was attached to the right achievement or employer. A true number in the wrong place is still a false claim.

Another gap is the skills section, where job-description vocabulary can migrate into a generated resume even when the person’s own history does not support that exact competency.

Those gaps are documented rather than hidden. I would much rather know exactly where a safeguard ends than call the system “hallucination free.” It is not.

The useful question was never whether a control sounds strong. It is what class of failure it can actually stop, and what still gets through.

07

WE DELIBERATELY TRIED TO BREAK THE SAFEGUARD

One thing I wanted to stop doing after the first project was treating a passing test as proof that a control was strong.

So the resume validator has a mutation test. Instead of feeding it good resumes and checking they pass, the test deliberately damages a known-good one. Remove a metric. Inflate a number. Insert a credential. Delete an employer. Move a real metric onto a different employer. Then see what the validator catches.

I particularly like the last one, because it tests a weakness we already know exists. There is something uncomfortable about building a test that proves your own safeguard is incomplete. That is exactly why it is useful. A red team that only demonstrates what you already know how to prevent is not really red teaming.

08

I TRIED TO SOLVE JUDGEMENT WITH REGEX. TWICE.

I found cases where terminology from a job description had leaked into a generated resume’s competency section. No achievement was invented, but the language made the experience sound closer to the role than the underlying history justified.

My first instinct was deterministic: write a checker, compare competency phrases against the career history, block anything unsupported. It sounded sensible. It failed. Words obviously related to a human reader looked different to a pattern matcher — “facilitation” and “facilitated,” “testing” and “tested.” Making the matching more flexible created the opposite problem: prefix matching accepted words that merely began similarly.

We tried again. The second approach failed differently.

The conclusion was uncomfortable and useful: this was not a string-matching problem. It was a judgement problem. A deterministic pass could shortlist suspicious phrases, but deciding whether two expressions represent the same underlying capability needs semantic judgement — not an increasingly elaborate regular expression pretending to understand meaning.

THE DISTINCTION THE WHOLE PROJECT TURNS ON

AN INSTRUCTION IS NOT
A CONTROL.

Both of these look like safety. Only one of them survives a model that is confident, fluent and wrong.

AN INSTRUCTIONWRITTEN DOWN
  1. 1The rule lives in a prompt or a document.
  2. 2Following it depends on the model’s judgement in the moment.
  3. 3A breach looks exactly like compliance until someone reads the output closely.
  4. 4Nothing stops the action from happening.
RESULT — HOPE, DOCUMENTED
A CONTROLBUILT IN
  1. 1The rule is enforced by code that runs after generation, or by a tool the process simply does not have.
  2. 2It returns a verdict — and can refuse outright.
  3. 3A breach becomes a blocked document rather than a delivered one.
  4. 4Its limits can be tested, measured and written down.
RESULT — A GATE THAT CAN SAY NO

The same reasoning produced the principle I now use everywhere: use deterministic logic where the answer can actually be determined, and use AI where judgement is genuinely required. Not the other way around.

09

THE SAME LESSON APPEARED IN JOB DISCOVERY

Job discovery sounds simple until real job boards are involved. Different platforms expose information differently, location data is inconsistent, some boards resist automation, and the same role may appear in several forms.

I ended up with adapters for several major recruitment systems and agencies, plus ingestion from job alerts and manually pasted links. I also made a deliberate decision not to fight sites that clearly did not want to be scraped. If a board blocked automated access, the answer was not to get cleverer at bypassing it — the sanctioned route was the site’s own job alerts. That trade was easy, because the objective was never “collect every possible job.” It was “build a search process that is useful and sustainable.”

Location was the other deceptively simple problem. A job might say AUS - NSW - MACQUARIE PARK. A human in Sydney knows what that means. A naive filter looking for the word “Sydney” throws it away.

So the cheap deterministic filter was made deliberately conservative: filter confidently at country level, leave the nuance to the later evaluation stage. On one measured sample across two real boards, 355 of 600 listings were removed with no false positives found among those drops. That does not prove the filter will never make a mistake. It shows why we were careful about what a cheap rule was allowed to decide.

10

SOMETIMES I CHOSE NOT TO FIX A BUG

One bug taught me something else. The job-description reader used whitespace to estimate whether a posting contained enough text. That worked in English and failed badly for languages such as Japanese.

My first instinct was to fix it. But a newer country filter already removed those jobs earlier, because they were outside the target geography. Fixing the reader first would have spent time and money processing work that another control would correctly discard.

Earlier, I would have treated every discovered bug as a task. Now I ask what consumes the output, what other controls already exist, and what the right sequence is. A technically correct fix can still be the wrong change.

11

AN AI TOLD ME SEVEN JOB BOARDS WERE DEAD. IT WAS WRONG.

This is one of my favourite mistakes, because I caught it and the AI did not.

A health report showed several boards had added zero new jobs. The AI helping me read it concluded the boards were not working. That sounded plausible. But the pipeline deduplicates jobs it has already seen. “Zero new” did not mean the board returned nothing — it could mean every job it returned that day was already known.

The metric was being interpreted without understanding the process that produced it. We changed the reporting so the dashboard explains why jobs were filtered or rejected, instead of showing only a final count.

The enterprise parallel is obvious. Dashboards are full of numbers that look self-explanatory: cases closed, contacts avoided, automation rate, AI containment, productivity. None of those numbers explains itself. You have to understand the process that produced it.

12

I DID NOT WANT ONE POWERFUL ASSISTANT

It would have been simpler to give one assistant access to everything and tell it to behave differently depending on the question. I chose not to.

Joblinko has several separate conversational doors. They look similar to the user and do not have the same permissions. A conversation about a job can read relevant information but cannot change files, run shell commands or reach the web. A career-history interview can write specific files because that is its purpose, without unrestricted system access. A troubleshooting conversation is prevented from quietly fixing the product while the user believes they are only asking for help.

Some of those restrictions live in code rather than in the prompt, and that difference matters. A prompt saying “do not change anything” is an instruction. A process that does not possess the tool required to change anything is a boundary.

The question for an enterprise copilot should not only be what the assistant should do. It should also be what that assistant is technically capable of doing if its reasoning goes wrong.

13

MORE AUTONOMY WAS NOT THE GOAL

The most obvious feature I chose not to build was automatic job applications. It would have been easy to tell a compelling story about volume: find more jobs, tailor automatically, fill the forms, submit everything, wake up to fifty applications.

A job application is a consequential representation of a person. The user should decide whether the opportunity is worth pursuing, review the resume, and decide whether something is submitted in their name. The system can prepare, recommend and remove repetitive work. It stops before the irreversible action.

That was a product decision, not an unfinished feature. Enterprise AI needs the same distinction: recommendation and authority are not the same thing.

14

I ALSO DECIDED NOT TO BUILD MY OWN VOICE TECHNOLOGY

Interview practice benefits from speaking rather than typing. I could have added speech recognition, text-to-speech, audio storage and another AI stack. I chose not to.

Phones and existing AI apps already handle voice well, so Joblinko prepares the grounded context, facts and rehearsal material and lets an existing voice system handle the conversation. No bespoke speech pipeline, and no stored audio anywhere.

It was less impressive as a feature list and better product judgement. Not every capability has to be rebuilt simply because it can be.

15

REAL USE CHANGED THE INTERVIEW PREPARATION

One of the most useful things about Joblinko is that another real person used it extensively, which immediately exposed assumptions I could not see.

The first deep preparation document was comprehensive and not very usable. It discussed an unfamiliar industry in that industry’s own language. Terms were technically correct and unexplained. The user told me she could not follow it.

That feedback changed the design. Industry context moved to the front. Technical terms had to be explained where they first appeared. Acronyms had to be expanded. Most importantly, anything the person might actually need to say in an interview had to be written as a complete sentence. “Mention policy administration modernisation” is useless under pressure. A sentence they understand and can naturally say is not.

It looks like a small content change. I think it reflects something bigger: AI frequently produces material for the expert who requested it, rather than the person who has to use it. The same thing happens in contact centres, where a knowledge article can be entirely correct and still be useless to an agent explaining something to an upset customer in twenty seconds.

16

THE SCORING SYSTEM TAUGHT ME TO DISTRUST NUMBERS

Joblinko evaluates opportunities across several dimensions — fit with experience, fit with target roles, compensation, potential concerns, and company or cultural considerations. The weights are configurable, and the system produces an overall score that makes comparison easier.

I do not want to pretend the number is more scientific than it is. Several of the underlying scores are model judgement. Even the final weighted calculation is currently produced by the model rather than independently recomputed by deterministic code.

More importantly, I have not demonstrated that a higher Joblinko score predicts a higher probability of an interview or an offer. There is no controlled outcome dataset behind that claim. So the score is a decision aid, not a prediction.

That distinction matters because a score changes behaviour whether or not anyone has proved it predicts anything. People sort by it, act on it and stop reading below it.

AI makes it very easy to produce numbers that look analytical. A number with decimal places is still an opinion if the mechanism behind it is judgement.

17

ONE OF MY BEST PROCESS IMPROVEMENTS ALSO FAILED

After the first project’s context problems, I introduced a clearer document hierarchy: permanent rules in one place, current state in another, agreed-but-unfinished decisions in a third. I also created a detailed handoff mechanism between AI sessions, so a new conversation would not have to reconstruct the project from scratch.

For a period it worked extremely well. The handoffs captured failures, decisions, measurements and explicit warnings not to repeat expensive investigations.

Then development accelerated, other work took priority, and the handoff stopped being maintained. The solution to context debt accumulated context debt.

I could describe that as a process failure. It was. But the more useful lesson is that a control relying on people remembering to maintain it will weaken precisely when the organisation gets busiest. The same is true of SOPs, knowledge bases, risk registers and exception logs. A good mechanism is not enough — the operating process has to keep it alive.

18

THE UNATTENDED RUN TAUGHT ME WHAT “DONE” CAN HIDE

At one point an overnight run queued around seventy jobs and evaluated none of them. The system still reported that the run had completed.

The AI session had expired, but the automation checked the wrong part of the command chain. The mechanics completed, so the system reported success even though the purpose of the run had failed.

That is the failure I worry about more than an obvious error. A red screen gets attention. A green screen attached to the wrong outcome creates confidence.

The eventual fix was deliberately cautious: fail loudly when the system knows the session is signed out, but never strand a healthy machine because an old client, a timeout or some ambiguous condition was mistaken for failure. “Fail safe” sounds simple until the safeguard itself can cause the outage.

EVIDENCE, AND ITS EDGES

WHAT IT PROVED.
AND WHAT IT DIDN’T.

The parts of this project I am most confident about are the parts where I can say precisely what the evidence does not reach.

BEYOND A JOB SEARCH

THE SAME PATTERNS,
IN CX AND BACK OFFICE.

The more I built, the less this felt like a job-search project. Every mechanism here has an obvious enterprise twin — and the failures translate even better than the features.

GROUNDED GENERATIONA resume grounded in a career history is not architecturally far from a customer communication grounded in an account record and a policy.SAME SHAPE
PRE-SEND CONTROLSA deterministic claim validator resembles a control checking whether an AI-generated claims letter contains statements the case actually supports.SAME SHAPE
BOUNDED COPILOTSA chat that can read but not act resembles an employee copilot that can explain a policy but cannot approve a refund.SAME SHAPE
APPROVAL POINTSA workflow that prepares everything and stops resembles a finance or claims assistant that recommends but requires a person to authorise the consequence.SAME SHAPE
KNOWLEDGE DECAYThe handoff document that quietly went stale looks exactly like an operations knowledge base that was excellent at launch and slowly stopped reflecting reality.SAME SHAPE
METRICS WITHOUT PROCESSThe misleading “zero new jobs” number looks like any operational dashboard whose figure cannot be interpreted without understanding the process beneath it.SAME SHAPE
THE HARDER PARTNOT WHERE TO PUT AI — WHAT KIND OF WORK EACH STEP ACTUALLY IS
WORK THAT IS DETERMINISTICAn authority limit, a mandatory disclosure or a regulatory threshold is the same kind of object as “every number in this document must appear in the source.” I built that rule as an instruction first, and the model went around it. I built it again as a gate that returns an exit code and refuses to produce the file, and it held. The difference was never the wording. It was whether anything could say no.BUILD THE GATE
WORK THAT NEEDS JUDGEMENTI spent two attempts trying to catch borrowed vocabulary with pattern matching, and both failed in both directions at once — false alarms on “facilitated” against “facilitation,” blind spots on a prefix that merely looked similar. A complaint with conflicting evidence, or an exception no rule anticipated, has that same shape. Forcing it into a rule does not produce a control. It produces something confidently wrong.LET IT JUDGE
WORK THAT NEEDS PERMISSION, NOT JUST CAPABILITYJoblinko finds the job, scores it, writes the application and prepares the interview. It cannot send anything, and that is not an unfinished feature. A model may be entirely capable of recommending a refund, flagging a transaction or drafting a decision, and still not be the thing that should execute it. Capability and authority are separate design decisions. Conflating them is how automation becomes an incident.ASK A HUMAN
THE OPERATING TRUTH, WHICH IS USUALLY SCATTEREDThis is the one I underestimated. Eleven facts that existed only in somebody’s memory were worth more than any model improvement — and in a real operation the equivalent is scattered much further: exception notes, client-specific instructions, supervisor judgement, contractual differences, the behaviour of a system nobody documented. Put AI on top of that and it inherits the conflicts rather than resolving them. If the exception path is undocumented, it cannot follow it.FIND IT FIRST
THE STARTING QUESTIONNOT “WHERE CAN WE PUT AI” — “WHAT IS THE OPERATING TRUTH, AND WHO IS ACCOUNTABLE WHEN IT IS WRONG”
WHICH PARTS ARE DETERMINISTIC · WHERE IS JUDGEMENT GENUINELY REQUIRED · WHAT MAY IT RECOMMEND · WHAT IS IT ALLOWED TO DO · WHERE MUST A PERSON STAY ACCOUNTABLE · WHAT EVIDENCE WOULD SHOW THE PROCESS IS BETTER RATHER THAN MERELY MORE AUTOMATED
THE MOST USEFUL ARCHITECTURE IS NOT THE ONE CONTAINING THE MOST AI. IT IS THE ONE THAT PUTS UNCERTAINTY ONLY WHERE UNCERTAINTY IS ACTUALLY NECESSARY.
WHAT I CAN HONESTLY CLAIM

I DECIDED WHAT
TO TRUST.

I did not create Joblinko from nothing. I started from a useful open-source foundation built by somebody else, then spent months adapting, extending, testing and changing it. AI wrote a large amount of the implementation. My role was to decide what problem we were solving, what had to stay true, what the AI should and should not be allowed to do, which proposed fixes were acceptable, what needed testing, what evidence was sufficient, and what should be removed or left unresolved. Along the way I got more comfortable saying: I don’t know. This has not been tested. This number is judgement, not prediction. This safeguard has a gap. This should not be built. I stopped treating those as admissions. The first project taught me how much sits between an impressive AI output and a dependable system. The second taught me the answer is not more control, more prompting or more AI — it is choosing the right mechanism for each part of the work, making the boundaries explicit, and letting the evidence tell you your design is wrong.

Which raised the next question. What changes when the same principles meet a process where the evidence is fragmented across hundreds of documents, privacy matters a great deal more, and an unsupported conclusion carries a real financial consequence?

SEE THE SYSTEM↗