Type Alias: LlamaModelOptions

type LlamaModelOptions = {
  modelPath: string;
  gpuLayers?:   | "auto"
     | "max"
     | number
     | {
     min?: number;
     max?: number;
     fitContext?: {
        contextSize?: number;
        embeddingContext?: boolean;
     };
   };
  vocabOnly?: boolean;
  useMmap?: boolean;
  useMlock?: boolean;
  checkTensors?: boolean;
  defaultContextFlashAttention?: boolean;
  defaultContextSwaFullCache?: boolean;
  onLoadProgress?: void;
  loadSignal?: AbortSignal;
  ignoreMemorySafetyChecks?: boolean;
  metadataOverrides?: OverridesObject<GgufMetadata, number | bigint | boolean | string>;
};

Defined in: evaluator/LlamaModel/LlamaModel.ts:26

Properties

modelPath

modelPath: string;

Defined in: evaluator/LlamaModel/LlamaModel.ts:28

path to the model on the filesystem

gpuLayers?

optional gpuLayers: 
  | "auto"
  | "max"
  | number
  | {
  min?: number;
  max?: number;
  fitContext?: {
     contextSize?: number;
     embeddingContext?: boolean;
  };
};

Defined in: evaluator/LlamaModel/LlamaModel.ts:44

Number of layers to store in VRAM.

"auto" - adapt to the current VRAM state and try to fit as many layers as possible in it. Takes into account the VRAM required to create a context with a contextSize set to "auto".
"max" - store all layers in VRAM. If there's not enough VRAM, an error will be thrown. Use with caution.
number - store the specified number of layers in VRAM. If there's not enough VRAM, an error will be thrown. Use with caution.
{min?: number, max?: number, fitContext?: {contextSize: number}} - adapt to the current VRAM state and try to fit as many layers as possible in it, but at least min and at most max layers. Set fitContext to the parameters of a context you intend to create with the model, so it'll take it into account in the calculations and leave enough memory for such a context.

If GPU support is disabled, will be set to 0 automatically.

Defaults to "auto".

vocabOnly?

optional vocabOnly: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:64

Only load the vocabulary, not weight tensors.

Useful when you only want to use the model to use its tokenizer but not for evaluation.

Defaults to false.

useMmap?

optional useMmap: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:77

Use mmap (memory-mapped file) to load the model.

Using mmap allows the OS to load the model tensors directly from the file on the filesystem, and makes it easier for the system to manage memory.

When using mmap, you might notice a delay the first time you actually use the model, which is caused by the OS itself loading the model into memory.

Defaults to true if the current system supports it.

useMlock?

optional useMlock: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:83

Force the system to keep the model in the RAM/VRAM. Use with caution as this can crash your system if the available resources are insufficient.

checkTensors?

optional checkTensors: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:91

Check for tensor validity before actually loading the model. Using it increases the time it takes to load the model.

Defaults to false.

defaultContextFlashAttention?

optional defaultContextFlashAttention: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:112

Enable flash attention by default for contexts created with this model. Only works with models that support flash attention.

Flash attention is an optimization in the attention mechanism that makes inference faster, more efficient and uses less memory.

The support for flash attention is currently experimental and may not always work as expected. Use with caution.

This option will be ignored if flash attention is not supported by the model.

Enabling this affects the calculations of default values for the model and contexts created with it as flash attention reduces the amount of memory required, which allows for more layers to be offloaded to the GPU and for context sizes to be bigger.

Defaults to false.

Upon flash attention exiting the experimental status, the default value will become true.

defaultContextSwaFullCache?

optional defaultContextSwaFullCache: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:123

When using SWA (Sliding Window Attention) on a supported model, extend the sliding window size to the current context size (meaning practically disabling SWA) by default for contexts created with this model.

See the swaFullCache option of the .createContext() method for more information.

Defaults to false.

loadSignal?

optional loadSignal: AbortSignal;

Defined in: evaluator/LlamaModel/LlamaModel.ts:132

An abort signal to abort the model load

ignoreMemorySafetyChecks?

optional ignoreMemorySafetyChecks: boolean;

Defined in: evaluator/LlamaModel/LlamaModel.ts:140

Ignore insufficient memory errors and continue with the model load. Can cause the process to crash if there's not enough VRAM to fit the model.

Defaults to false.

metadataOverrides?

optional metadataOverrides: OverridesObject<GgufMetadata, number | bigint | boolean | string>;

Defined in: evaluator/LlamaModel/LlamaModel.ts:149

Metadata overrides to load the model with.

Note: Most metadata value overrides aren't supported and overriding them will have no effect on llama.cpp. Only use this for metadata values that are explicitly documented to be supported by llama.cpp to be overridden, and only in cases when this is crucial, as this is not guaranteed to always work as expected.

Methods

onLoadProgress()?

optional onLoadProgress(loadProgress: number): void;

Defined in: evaluator/LlamaModel/LlamaModel.ts:129

Called with the load percentage when the model is being loaded.

Parameters

Parameter	Type	Description
`loadProgress`	`number`	a number between 0 (exclusive) and 1 (inclusive).

Returns

void

LlamaModel

LlamaModelTokens

LlamaChatSession

LlamaText

GgufInsights

GbnfJsonSchema

ChatHistoryItem

ChatModelResponse

LlamaChatResponse

GgufFileInfo

GgufMetadata

LlamaContextOptions

BatchingOptions

LlamaChatSessionOptions

LLamaChatPromptOptions

Chat Wrapper Options

JinjaTemplateChatWrapperOptions

Type Alias: LlamaModelOptions

Properties

modelPath

gpuLayers?

vocabOnly?

useMmap?

useMlock?

checkTensors?

defaultContextFlashAttention?

defaultContextSwaFullCache?

loadSignal?

ignoreMemorySafetyChecks?

metadataOverrides?

Methods

onLoadProgress()?

Parameters

Returns

LlamaModelTokens

ChatModelResponse

GgufMetadata

LlamaContextOptions

BatchingOptions

LlamaChatSessionOptions

LLamaChatPromptOptions

JinjaTemplateChatWrapperOptions

Type Alias: LlamaModelOptions ​

Properties ​

modelPath ​

gpuLayers? ​

vocabOnly? ​

useMmap? ​

useMlock? ​

checkTensors? ​

defaultContextFlashAttention? ​

defaultContextSwaFullCache? ​

loadSignal? ​

ignoreMemorySafetyChecks? ​

metadataOverrides? ​

Methods ​

onLoadProgress()? ​

Parameters ​

Returns ​

Type Alias: LlamaModelOptions

Properties

modelPath

gpuLayers?

vocabOnly?

useMmap?

useMlock?

checkTensors?

defaultContextFlashAttention?

defaultContextSwaFullCache?

loadSignal?

ignoreMemorySafetyChecks?

metadataOverrides?

Methods

onLoadProgress()?

Parameters

Returns