Jelia.nyc

Jelia.nyc / What I do / AI that earns its keep

WHAT I DO ยท AI

AI where it removes work, and nowhere else

I find the places where a model genuinely takes work off your people, and I say plainly where it will not. Then I build it into the systems you already use, with costs you can predict.

What it is

Practical AI work, not a strategy document. Find a task your people do repeatedly, decide whether a model can do it well enough, prove that against your own examples, and then wire it into the place where the work actually happens. If the answer is that a model will not help, you hear that too. It is one of the more useful things I can tell you.

Where it earns its keep, and where it does not

Models are good at:

  • Drafting and summarising, where a person reviews the result before it goes anywhere.
  • Sorting, tagging and routing incoming material: tickets, enquiries, documents.
  • Answering questions from your own documentation, with the sources shown.
  • Engineering and operations work under review: scripts, configuration, investigation, write-ups.

They are poor at, or the wrong tool for:

  • Decisions that someone has to be accountable for.
  • Figures of record. A model is not a calculator or a ledger.
  • Anything where a wrong answer is expensive and nobody would notice it was wrong.
  • Covering for a broken process. Fix the process first. Sometimes that is the whole job.

How I do it

  1. Find the work. Which task, how often, how long it takes now, and what a mistake costs.
  2. Build an evaluation set. Real examples from your own work with known good answers. This is how we will judge any model, before anyone relies on it.
  3. Choose the model and the endpoint. Whichever performs best on your examples at a sensible cost. Private or business endpoints whose data terms suit what you are sending, checked with you rather than assumed.
  4. Integrate properly. Into your ticketing, mail, documents or code, not a separate chat window people have to remember to open.
  5. Account for cost per job. Every call is attributed to the task that made it, with limits, so the bill is predictable and a runaway loop is caught.
  6. Keep a human where it matters. Review steps sit where a mistake would be costly, and nowhere else.
  7. Roll out, then re-test. The evaluation set runs again whenever the model, the prompt or the data changes.

What you get

  • A short written assessment: where AI will help in your organisation, where it will not, and why.
  • Working integrations for the tasks that pass evaluation.
  • The evaluation set and its results, so you can re-test after any change.
  • Per-job cost reporting and spending limits.
  • Documentation your team can maintain without me.

In my own estate

Claude Code does the day-to-day engineering and operations work across my infrastructure. The Anthropic API powers an Ask Claude widget across 25 sites and the scheduled agents that keep the newsroom current, with per-job cost accounting on all of it. GitHub Copilot and OpenAI's GPT models are kept deliberately in the mix, so that no single vendor holds the estate and each job goes to whichever model suits it.

25Sites with the Ask Claude widget
Per jobCost accounting on every model call
More than oneModel vendor, by design

Questions people ask

Will this replace staff?

That is not how I approach it. The aim is to take the repetitive part of a job off a person so they spend the time on the part that needs them.

Will our data be used to train someone's model?

That depends on the provider and the terms of the account you use. I check those terms with you, and pick endpoints whose terms suit the data you will be sending. Some data should not go to an external model at all, and I will say so.

Which model do you recommend?

Whichever does best on your evaluation set at a cost you are comfortable with. I use several vendors myself, deliberately.