AI / 03 Urquhart / capability

Local AI computing

Plan local AI around models, memory and operating reality.

Architecture for local inference and AI development systems, with attention to VRAM, RAM, storage, cooling, power, privacy and the upgrade path.

Who it is for

A defined service for a recognisable technical problem.

  • Developers experimenting with local models
  • Businesses exploring private inference
  • Researchers who need a bounded proof-of-concept environment

Problems addressed

  • Model ambitions that exceed available VRAM or memory
  • Poorly balanced GPU, CPU, storage and power choices
  • Unclear privacy, access and software-stack requirements
  • Buying expensive hardware before proving the workload

Included

  • Use-case, model and concurrency assumptions
  • Hardware architecture and constraint analysis
  • Software-stack recommendation
  • A staged route from feasibility to deployment

Not included

  • Model performance or commercial outcome guarantees
  • Training large foundation models from scratch
  • Third-party API, model, licence and hardware costs
  • Production deployment without separate security and operational scope

Delivery process

Reduce uncertainty in deliberate stages.

01

Bound the use case

Define models, context, throughput, users, privacy and budget.

02

Test assumptions

Identify memory, compute and software constraints before committing.

03

Specify

Produce a balanced architecture and staged purchasing route.

04

Configure

Optional runtime setup and initial validation follow the approved architecture.

Starting prices

Choose the smallest useful commitment.

Models, APIs, hosted services, licences and hardware are separate third-party costs.

See the complete price list

AI or Local-LLM Workstation Architecture

Model, VRAM, RAM, storage, cooling and upgrade path

£495

AI Workstation Software Setup

Local AI runtime and initial environment

From £395

Local AI Feasibility Review

Hardware, model, privacy and workload

£795

Evidence and limits

Credibility comes from what can be checked.

  • Architecture is tied to named model classes and explicit workload assumptions.
  • Memory limits, quantisation choices and operational trade-offs are documented.
  • A feasibility review can precede hardware purchase when uncertainty is material.

Questions

Before you enquire.

Do I need the most expensive GPU?

Not necessarily. The right choice depends on model size, precision, concurrency, latency and budget.

Can local AI keep data private?

Local processing can reduce external data transfer, but privacy still depends on access controls, logging, backups and the wider system design.

Can you guarantee a model will meet my needs?

No. The correct route is to define an evaluation and test representative material before treating a model as production-ready.

Start with evidence

Describe the current system and the outcome you need.

Include the main constraint, the people affected, the timescale and any existing specification. UDS will confirm the most proportionate first step.

Send the project brief