Compute
Operations, execution resources, and the useful work performed during inference.
The computational frontier
Frontier research in AI infrastructure and computational efficiency.
We investigate how intelligence is computed—and how its foundations can make it more efficient, affordable, and accessible.
Capability alone does not make intelligence accessible. The resources needed to execute it determine who can use it—and at what cost.
Memory capacity, bandwidth, computational throughput, execution scheduling, and hardware constraints shape the economics of intelligent systems. We treat these constraints as research opportunities.
Forera is an independent research-driven engineering venture. We study how computational work is organized, measured, and executed, seeking deeper understanding that can inform practical improvements in AI infrastructure.
Our ambition is to change the computational economics of intelligent systems. It is a direction for investigation, not a claim of a breakthrough.
Six dimensions. One interdependent system.
An improvement in one place can move the cost elsewhere.
Operations, execution resources, and the useful work performed during inference.
The space for model state, and the rate at which that state can be accessed.
Transfers between storage, memory, and compute can dominate execution time.
How work is ordered, overlapped, and distributed across available resources.
The numerical formats and structures used for weights and intermediate states.
Instrumentation that separates a real bottleneck from an attractive assumption.
Precision is a tradeoff. Lower numerical precision can reduce memory demand while changing accuracy or hardware efficiency.
Concurrency has a cost. Overlapping more work can improve throughput while increasing memory pressure or latency.
Our current research interests.
Open directions, investigated through experiments and engineering.
Investigating how model execution is structured, and where alternative architectures may change performance, resource needs, and scale.
How should model execution be organized?
Studying useful computational work in relation to hardware utilization, memory behavior, and the cost of execution.
Where does useful work become overhead?
Developing instrumentation, profiling methods, and reproducible environments for controlled computational experiments.
What evidence makes a comparison meaningful?
Exploring infrastructure for composing, coordinating, and observing intelligent computational workflows.
How do complex processes remain accountable?
Investigating how systems adapt to hardware constraints and make better use of the resources already available.
What can intelligence do within a budget?
Exploring hypotheses and architectural alternatives that challenge assumptions in established AI execution systems.
Which assumptions deserve another experiment?
Research proceeds in loops, not declarations. Every explanation is open to revision by the next experiment.
Clear baselines. Disclosed hardware and software configurations. Controlled variables. Correctness and quality checks. Reproducible comparisons. Transparent limitations.
Instrument behavior. Identify what can be measured.
Explain mechanisms. Form a testable hypothesis.
Isolate variables. Compare controlled alternatives.
Check correctness, quality, and reproducibility.
Build reusable infrastructure from validated findings.
Let new measurements revise the explanation.
New measurements return to observation.
The initial Forera ecosystem.
Experimental projects that connect questions to working systems.
Computational research / AI inference
An experimental project investigating how large language models execute. Cachalot examines inference mechanisms, resource consumption, and opportunities for alternative execution strategies.
Instrumentation / Experimental systems
The macOS observability and experimentation environment associated with Cachalot. Lab makes runtime activity visible; its research direction includes deeper profiling, structured comparisons, and reproducible experiments.
Intelligent systems / Agent orchestration
Crewshal explores reusable, project-independent infrastructure for coordinating intelligent workflows and computational agents. The focus is modular execution, adaptable architecture, and reliable coordination.
Faster compute is only useful when data can keep up. Explore the relationship between arithmetic work and memory bandwidth.
P ≤ min(Ppeak, B × I)Change the workload.
Watch the constraint.
More arithmetic per byte or more memory bandwidth raises this bound. The compute ceiling still applies.
A simplified upper bound, not a benchmark or a prediction for Cachalot. Real performance also depends on latency, scheduling, cache behavior, numerical formats, and correctness. Units use decimal GB and GFLOP.
Before the ridge, the memory system limits the bound: bandwidth multiplied by operations per byte. Beyond it, peak compute throughput sets the ceiling. Raising one resource does not remove every constraint.
Roofline model · Berkeley LabRigor is a practice.
Not an aesthetic.
Investigate the mechanism before changing its behavior.
Claims need reproducible experiments and clearly defined baselines.
How work is organized matters as much as the speed of individual operations.
Evaluate performance alongside correctness, quality, reliability, and practical constraints.
Seek knowledge that can inform real systems and usable infrastructure.
Established techniques are foundations for investigation, not its boundaries.
Our research starts with questions about computational behavior and architectural efficiency. Promising findings may become experimental implementations, open-source tools, reusable libraries, infrastructure components, or developer platforms.
Commercial technologies and technical collaborations can put those ideas to work—and expose new questions for research. Scientific investigation and practical engineering reinforce one another.
A conceptual path, not a record of completed stages. Progress depends on evidence.
For researchers, engineers, and organizations.
We welcome meaningful technical conversations about advancing the efficiency and accessibility of intelligent computation.
Start a conversation