ollama

mirror of https://github.com/ollama/ollama synced 2026-04-23 08:45:14 +00:00

Author	SHA1	Message	Date
Parth Sareen	123b300af6	docs: update hermes (#15655 )	2026-04-17 14:20:59 -07:00
Parth Sareen	57653b8e42	cmd/launch: show WSL guidance on Windows instead of handing off (#15637 )	2026-04-16 17:18:04 -07:00
Parth Sareen	a50ce61c54	launch: skip unchanged managed-single rewrite (#15633 )	2026-04-16 16:20:42 -07:00
Daniel Hiltgen	2bb7ea00d2	create: avoid gc race with create (#15628 ) If you have a long running create, and start another ollama server with the same model dir, the GC algorithm deletes the pending blobs and breaks the create. This adds a 1h grace period to avoid deleting in-flight creation operations.	2026-04-16 13:29:16 -07:00
Daniel Hiltgen	55fa80d07a	mlx: additional gemma4 cache fixes (#15607 ) Harden additional corner cases	2026-04-16 13:07:19 -07:00
Daniel Hiltgen	b9cb535407	mlx: fix gemma4 cache to use logical view (#15617 )	2026-04-16 11:54:30 -07:00
Daniel Hiltgen	031baef094	mlx: fix imagegen lookup (#15588 ) * mlx: fix imagegen lookup Fixes #15533 - imagegen had fallen out of sync with the new layout for multiple mlx libraries on Metal. * review comments	2026-04-16 10:39:00 -07:00
Mike Wallio	7d271e6dc9	cmd/launch: add Copilot CLI integration (#15583 ) --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2026-04-15 17:22:53 -07:00
Devon Rifkin	c88dae2d6b	Merge pull request #15612 from ollama/drifkin/gemma4-split-templates gemma4: render differently based on model size	2026-04-15 17:15:35 -07:00
Devon Rifkin	9e3618d663	make empty block conditional	2026-04-15 15:35:25 -07:00
Daniel Hiltgen	5d920cc6bc	Keep Gemma4 router projection in source precision (#15613 )	2026-04-15 15:04:23 -07:00
Devon Rifkin	e585ecd11f	gemma4: render differently based on model size Following up on #15560, this change now has e2b/e4b render differently from 26b/31b. For backwards compatibility, we take the existing renderer name `gemma4` and make it do dynamic resolution based on the model name/size, but the intended use is for the models to be republished with the renderer variant specified explicitly: `gemma4-small` or `gemma4-large`.	2026-04-15 14:37:16 -07:00
Eva H	cdddea0592	launch: always list cloud recommendations first (#15593 )	2026-04-15 13:17:35 -07:00
Parth Sareen	43f90def04	launch: add hermes (#15569 )	2026-04-15 12:00:23 -07:00
Daniel Hiltgen	06ae6367bd	mlx: fix RotatingKVCache.concat() dropping context on mid-rotation (#15591 ) After the rotating buffer has wrapped (c.offset > c.maxSize) a subsequent L>1 Update() went through a slice-to-[0, c.idx) path that discarded all slots in [c.idx, Dim), losing the older-but-still-in-window tokens the first Q of the new batch needs for its sliding-window attention. Linearize the circular buffer to logical order in that wrapped case so the existing trim + concat preserves the last (maxSize - 1) old tokens. When the buffer has not yet wrapped (c.offset <= c.maxSize), slots [c.idx, Dim) are grow padding or stale post-rewind data, so keep dropping them.	2026-04-14 18:29:06 -07:00
Daniel Hiltgen	48ad7085c4	mlx: Improve gemma4 performance with fused operations (#15587 ) * mlx: Improve gemma4 performance with fused operations * review comments	2026-04-14 18:04:04 -07:00
Jesse Gross	e1e3cec8d0	models: fuse MLP activation functions via mlx_compile Converts SiLU/GELUApprox to compiled kernels and adds SwiGLU, matching upstream mlx/mlx_lm's activations pattern. Routes llama, qwen3, qwen3_5 (dense + MoE), and glm4_moe_lite MLP paths through mlx.SwiGLU so each MLP invocation runs as one fused Metal/CUDA kernel rather than a chain of per-op launches.	2026-04-14 16:38:32 -07:00
Jesse Gross	d3e67e305c	mlx: add compiled closure support Wraps MLX's mlx_compile API so Go functions can be traced into fused kernels. Contiguous elementwise chains collapse into a single Metal/CUDA kernel instead of launching one per op. Exposes Compile plus arity helpers (Compile1/2/3) that mirror Python's @mx.compile decorator shape, lazily building the closure on first call so package-level declarations work before the MLX dylib loads.	2026-04-14 16:38:32 -07:00
Eva H	698e04a14b	launch: OpenCode inline config (#15586 )	2026-04-14 15:08:42 -07:00
Eva H	1d9537bc33	launch/openclaw: fix --yes flag behaviour to skip channels configuration (#15589 )	2026-04-14 13:57:35 -07:00
Eva H	120424d832	Revert "launch/opencode: use inline config (#15462 )" (#15568 )	2026-04-13 18:40:17 -07:00
Eva H	5818001610	launch: skip unchanged integration rewrite configration (#15491 )	2026-04-13 17:18:56 -07:00
Daniel Hiltgen	2cba7756c5	Gemma4 on MLX (#15244 ) * gemma4: implement Gemma 4 model for MLX (text-only runtime) * gemma4: two MoE + SWA prefill perf fixes Two performance optimizations in the gemma4 forward pass 1. Memoize the sliding-window prefill mask across layers. 2. Softmax only over the selected experts in Router.Forward. * review comments	2026-04-13 16:36:51 -07:00
Devon Rifkin	bf2a421727	gemma4: restore e2b-style nothink prompt (#15560 ) Gemma 4 prompts differ when thinking is disabled for different sized models: 26b/31b emit an empty thought block, while e2b/e4b do not. Before #15490, our shared Gemma 4 renderer effectively matched the e2b behavior. #15490 changed it to always emit the empty thought block, which regressed e2b/e4b nothink behavior and led to #15536 (and possibly This change restores the previous shared behavior by removing the empty trailing thought block. It also renames the checked-in upstream chat templates so the e2b and 31b fixtures are tracked separately. A follow-up will split Gemma 4 rendering by model size. Fixes: #15536	2026-04-13 14:26:15 -07:00
Eva H	f3cf6b75fb	launch/opencode: use inline config (#15462 )	2026-04-13 13:41:31 -07:00
Devon Rifkin	5dfac387a6	Revert "gemma4: fix nothink case renderer (#15553 )" (#15556 ) This reverts commit `4d75f5da03`.	2026-04-13 13:12:18 -07:00
Daniel Hiltgen	a99e5d9c22	mac: prevent generate on cross-compiles (#15120 ) For some versions of Xcode, cmake builds are failing due to header problems in cross-compiling during the generate phase. Since generate is producing arch independent generated output, we can skip this during cross-compiling.	2026-04-13 13:04:58 -07:00
Daniel Hiltgen	0abf3aca36	cgo: suppress deprecated warning to quiet down go build (#15438 )	2026-04-13 13:04:11 -07:00
Devon Rifkin	ee0266462a	Revert "gemma4: add nothink renderer tests (#15554 )" (#15555 ) This reverts commit `1b70bb8a10`.	2026-04-13 13:00:59 -07:00
Daniel Hiltgen	c88fb286ec	mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA (#14913 ) * mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA Add Conv2d, flexible Pad (with axes/mode), PadConstant, Maximum, Minimum, Softplus, ReLU, GLU, Clamp, Sin, Cos, Clip, ScaledDotProductAttentionMasked, and RoPEWithFreqs. Refactor RoPEWithBase to delegate to RoPEWithFreqs. * review comments * mlx: fix ScaledDotProductAttentionMasked to consult the mask argument	2026-04-13 11:43:24 -07:00
Daniel Hiltgen	d3da29cbfc	mlx: mixed-precision quant and capability detection improvements (#15409 ) Improve the MLX model creation pipeline with several model-agnostic changes: - Rewrite supportsVision to use vision_config instead of architecture name - Add supportsAudio for audio encoder detection - Add alignment checking (isAligned) for quantization group sizes - Support per-projection mixed quantization in MoE expert packing - Record per-tensor quant metadata in safetensors blobs - Parse per-tensor quant metadata at model load time - Validate quantize output is non-empty before storing - Fix pin/unpin cleanup in expert group quantization - Promote v_proj/k_proj/down_proj to INT8 for INT4 base quant - Add MetalIsAvailable() utility - Skip audio encoder tensors from quantization	2026-04-13 11:43:07 -07:00
Devon Rifkin	1b70bb8a10	gemma4: add nothink renderer tests (#15554 ) Meant to include in #15553	2026-04-13 11:38:19 -07:00
Daniel Hiltgen	ec29ce4ce3	gemma4: fix compiler error on metal (#15550 ) On some systems, the metal runtime compiler is failing due to an uninitialized variable from #15378. Fixes #15548	2026-04-13 11:32:00 -07:00
Devon Rifkin	4d75f5da03	gemma4: fix nothink case renderer (#15553 ) Regressed in #15490 Fixes: #15536	2026-04-13 11:23:19 -07:00
saman-amd	798fd09bfe	Update to ROCm 7.2.1 (#15483 ) Co-authored-by: Samiii777 <58442200+Samiii777@users.noreply.github.com>	2026-04-12 12:11:58 -07:00
Devon Rifkin	9330bb9120	gemma4: be less strict about whitespace before bare keys (#15494 )	2026-04-11 16:30:27 -07:00
Devon Rifkin	40a1317dfd	gemma4: update renderer to match new jinja template (#15490 ) * gemma4: update renderer to match new jinja template Google has updated their jinja template for gemma4, and so this change gives us parity with the new template. The parsing also slightly changed upstream, so we make a small change to our parser as well. I've also corrected a few probably existing edge cases, especially around type unions. The upstream output format is weird (a stringified array), but in practice the models seem to understand it well. * gemma4: special case simple `AnyOf`s The upstream template doesn't handle `AnyOf`s, but since in the previous commit we saw type unions work reasonably well, I'm now treating very simple `AnyOf`s as type unions to help in cases where they might be used * fix lint * gemma4: prefer empty instead of `None` We can't currently distinguish between a result being not-present vs. empty. The empty case seems more important (e.g., a legitimately empty tool call) * gemma4: be more careful for tool results with missing IDs	2026-04-10 15:45:27 -07:00
Devon Rifkin	fdfe9cec98	model/parsers: fix missing parallel tool call indices (#15467 ) We were missing setting the function index for several models that can make parallel tool calls. In the future we may want to consider putting some sort of post-parse hook and relieve the parsers of this duty. Fixes: #15457	2026-04-10 15:23:21 -07:00
Matteo Celani	9517864603	app/ui: re-validate image attachments when selected model changes (#15272 )	2026-04-10 14:03:51 -07:00
Bruce MacDonald	8e6d86dbe3	docs: add hermes agent integration guide (#15488 ) Update cloud and local model recommendations to match current models.go: add qwen3.5:cloud and glm-5.1:cloud, replace glm-4.7-flash with gemma4 and qwen3.5 as local options. Add documentation for Hermes Agent by Nous Research, covering installation, Ollama setup via custom endpoint, messaging configuration, and recommended models.	2026-04-10 13:13:36 -07:00
Parth Sareen	80d3744c5d	launch: update openclaw channel message (#15463 )	2026-04-09 15:20:30 -07:00
Eva H	2a94f03823	launch: add re-run hint to dependency error message (#15439 )	2026-04-09 09:51:34 -07:00
Patrick Devine	eb97274e5c	modelfiles: fix /save command and add shortname for safetensors based models (#15413 ) This change fixes two issues with Modelfiles: 1. If a user uses `ollama show --modelfile` to show a safetensors based model, the Model would leave the "FROM" field blank which won't allow a user to recreate the model. This change adds the model's current canonical short name to the FROM field. 2. If a user uses the `/save` command in the CLI any messages which were saved in a previous model wouldn't get saved (only the set of messages from the current session).	2026-04-08 21:05:39 -07:00
Daniel Hiltgen	6b5db12aa2	mlx: remove stale x86 libmlx library (#15443 ) Fixes #15433	2026-04-08 20:51:47 -07:00
7. Sun	612f0a17d3	fix: improve error message for unknown input item type in responses API (#15424 ) The default branch in unmarshalResponsesInputItem had two issues: - It referenced typeField.Type instead of itemType; these differ when the shorthand role-based format promotes an empty type to "message", meaning an unhandled type would show the wrong value in the error string. - It used %s formatting, so an empty type field produced the unhelpful message "unknown input item type: " with no indication what was missing. Fix by using itemType (the resolved value) with %q quoting, and add a dedicated message when itemType is empty (both type and role absent): "input item missing required 'type' field". Tests added for the empty-type and missing-type cases. Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-08 17:41:12 -07:00
Parth Sareen	673726fa0e	app: restore launch default and refine launch sidebar open for app (#15437 )	2026-04-08 16:59:21 -07:00
Daniel Hiltgen	b5918f9785	pull/push: refine safetensors (#14946 ) * pull: refine safetensors pull - Body drain in resolve() — drain response body before close so Go's HTTP client can reuse TCP connections instead of opening a new one per blob (1,075 extra TCP+TLS handshakes eliminated) - Skip speed recording for tiny blobs (<100KB) — prevents HTTP-overhead-dominated transfer times from poisoning the median, which the stall detector uses to cancel "too slow" downloads - Resume support for large blobs (>=64MB) — on failure, preserves partial .tmp files; on retry, re-hashes existing datak and sends Range header to download only remaining bytes; gracefully falls back to full download if server returns 200 instead of 206; SHA256 verification catches corrupt partials * harden push - Prevents killing TCP connections after every request. - Stronger backoff to handle server back-pressure and rate limiting - Larger buffered reads for improve safetensor upload performance - Better error message handling from server - Handle 201 if server says blob exists - Fix progress reporting on already uploaded blobs - Trace logging to help troubleshoot and tune going forward * review comments * review comments	2026-04-08 14:15:39 -07:00
Eva H	d17f482d50	launch/opencode: detect `curl` installed opencode at `~/.opencode/bin` (#15197 )	2026-04-08 13:54:51 -07:00
Parth Sareen	4e16f562c0	launch: add openclaw channels setup (#15407 )	2026-04-08 13:25:27 -07:00
Parth Sareen	55308f1421	launch: update ctx length for glm-5.1 and gemma4 (#15411 ) Also adds glm-5.1 in recommended models	2026-04-08 12:11:50 -07:00

1 2 3 4 5 ...

5325 commits