Skip to content

Class: LlamaContext ​

Defined in: evaluator/LlamaContext/LlamaContext.ts:74

Properties ​

onDispose ​

ts
readonly onDispose: EventRelay<void>;

Defined in: evaluator/LlamaContext/LlamaContext.ts:110

Accessors ​

disposed ​

Get Signature ​

ts
get disposed(): boolean;

Defined in: evaluator/LlamaContext/LlamaContext.ts:233

Returns ​

boolean


model ​

Get Signature ​

ts
get model(): LlamaModel;

Defined in: evaluator/LlamaContext/LlamaContext.ts:237

Returns ​

LlamaModel


contextSize ​

Get Signature ​

ts
get contextSize(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:241

Returns ​

number


batchSize ​

Get Signature ​

ts
get batchSize(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:245

Returns ​

number


flashAttention ​

Get Signature ​

ts
get flashAttention(): boolean | "auto";

Defined in: evaluator/LlamaContext/LlamaContext.ts:249

Returns ​

boolean | "auto"


kvCacheKeyType ​

Get Signature ​

ts
get kvCacheKeyType(): GgmlType;

Defined in: evaluator/LlamaContext/LlamaContext.ts:253

Returns ​

GgmlType


kvCacheValueType ​

Get Signature ​

ts
get kvCacheValueType(): GgmlType;

Defined in: evaluator/LlamaContext/LlamaContext.ts:257

Returns ​

GgmlType


stateSize ​

Get Signature ​

ts
get stateSize(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:265

The actual size of the state in the memory in bytes. This value is provided by llama.cpp and doesn't include all the memory overhead of the context.

Returns ​

number


currentThreads ​

Get Signature ​

ts
get currentThreads(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:272

The number of threads currently used to evaluate tokens

Returns ​

number


idealThreads ​

Get Signature ​

ts
get idealThreads(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:283

The number of threads that are preferred to be used to evaluate tokens.

The actual number of threads used may be lower when other evaluations are running in parallel.

Returns ​

number


totalSequences ​

Get Signature ​

ts
get totalSequences(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:296

Returns ​

number


sequencesLeft ​

Get Signature ​

ts
get sequencesLeft(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:300

Returns ​

number


memoryUsage ​

Get Signature ​

ts
get memoryUsage(): {
  ram: number;
  vram: number;
};

Defined in: evaluator/LlamaContext/LlamaContext.ts:305

Assumed memory footprint of the context in bytes

Returns ​
ts
{
  ram: number;
  vram: number;
}
ram ​
ts
ram: number;
vram ​
ts
vram: number;

Methods ​

dispose() ​

ts
dispose(): Promise<void>;

Defined in: evaluator/LlamaContext/LlamaContext.ts:219

Returns ​

Promise<void>


getAllocatedContextSize() ​

ts
getAllocatedContextSize(): number;

Defined in: evaluator/LlamaContext/LlamaContext.ts:287

Returns ​

number


getSequence() ​

ts
getSequence(options?: {
  contextShift?: ContextShiftOptions;
  tokenPredictor?: TokenPredictor;
  checkpoints?: {
     max?: number;
     interval?: number | false;
     maxMemory?: number | null;
  };
}): LlamaContextSequence;

Defined in: evaluator/LlamaContext/LlamaContext.ts:319

Before calling this method, make sure to call sequencesLeft to check if there are any sequences left. When there are no sequences left, this method will throw an error.

Parameters ​

ParameterTypeDescription
options{ contextShift?: ContextShiftOptions; tokenPredictor?: TokenPredictor; checkpoints?: { max?: number; interval?: number | false; maxMemory?: number | null; }; }-
options.contextShift?ContextShiftOptions-
options.tokenPredictor?TokenPredictorToken predictor to use for the sequence. Don't share the same token predictor between multiple sequences. Using a token predictor doesn't affect the generation output itself - it only allows for greater parallelization of the token evaluation to speed up the generation. > Note: that if a token predictor is too resource intensive, > it can slow down the generation process due to the overhead of running the predictor. > > Testing the effectiveness of a token predictor on the target machine is recommended before using it in production. Automatically disposed when disposing the sequence. See Using Token Predictors
options.checkpoints?{ max?: number; interval?: number | false; maxMemory?: number | null; }The maximum number of checkpoint to keep for the sequence when needed. When reusing a prefix evaluation state is not possible for the context sequence (like in contexts from recurrent and hybrid models, or with models that use SWA (Sliding Window Attention) when the swaFullCache option is not enabled on the context), storing checkpoints allows reusing the context state at certain points in the sequence to speed up the evaluation when erasing parts of the context state that come after those points. Those checkpoints will automatically be used when trying to erase parts of the context state that come after a checkpointed state, and be freed from memory when no longer relevant. Those checkpoints are relatively lightweight compared to saving the entire state, but taking too many checkpoints can increase memory usage. Checkpoints are stored in the RAM (not VRAM). See LlamaContextSequence.takeCheckpoint for more details on how checkpoints are taken and used.
options.checkpoints.max?numberThe maximum number of checkpoints to keep for the sequence when needed. Defaults to 32.
options.checkpoints.interval?number | falseTake a checkpoint every interval tokens when the sequence needs taking checkpoints. Defaults to 8192.
options.checkpoints.maxMemory?number | nullThe maximum memory in bytes to use for checkpoints for the sequence when needed. When taking a checkpoint causes the checkpoints pool memory to exceed this value, older checkpoints will be pruned until the total checkpoints memory usage is under this limit, while ensuring that at least one checkpoint is kept. Defaults to null (no memory limit).

Returns ​

LlamaContextSequence


dispatchPendingBatch() ​

ts
dispatchPendingBatch(): void;

Defined in: evaluator/LlamaContext/LlamaContext.ts:416

Returns ​

void


printTimings() ​

ts
printTimings(): Promise<void>;

Defined in: evaluator/LlamaContext/LlamaContext.ts:733

Print the timings of token evaluation since that last print for this context.

Requires the performanceTracking option to be enabled.

Note: it prints on the LlamaLogLevel.info level, so if you set the level of your Llama instance higher than that, it won't print anything.

Returns ​

Promise<void>