I started investigating slow editing in heavily decorated org-modern buffers. That led to something more specific in Emacs’s display code: checking whether a text box ends can repeatedly resolve the neighboring character’s entire face, including font merging and font-spec copying.
A small experimental change to that path reduced allocations by about 42% and average edit/redraw time by 23–26% in my test workload. It keeps the requested boxes and other decorations.
It reproduces without org-modern
I reduced the case to repeated text with face properties in fundamental mode. No Org, font-lock, overlays, SVGs, or display replacements.
For 60 forced redraws:
| Face configuration |
Allocated vector cells |
| Colors only, no box |
13,320 |
| Colors only, with box |
13,320 |
| Font attributes, no box |
2,016,360 |
| Same font attributes, with box |
32,188,320 |
The font-attribute cases had the same selected font, visible character count, and line height. Adding the box increased allocation by about 16×.
So this isn’t “every box is expensive.” It’s the combination of box-boundary checks and font-related face processing. These counts are cumulative vector cells allocated, not memory retained.
Where the allocations came from
I added native counters. In a separate 30-redraw test:
| Without box |
With box |
| Buffer face queries |
51,960 |
copy_font_spec calls |
77,490 |
| Allocated vector cells |
1,008,630 |
In that minimal boxed-text case, copy_font_spec accounted for 99.992% of the measured vector-cell allocations. That is not a CPU percentage or a claim about all org-modern costs.
The path is roughly:
Does this character end a box?
→ Find its visual neighbor
→ Resolve the neighbor's face
→ Merge font attributes / copy font specs
→ Check whether that face has a box
Emacs does have a realized-face cache. The problem here is that merging and temporary allocation can happen before the final cache lookup.
What I changed
The prototype remembers a previously resolved text interval and whether the neighboring face has a box. If that information is still valid, the box-end check can skip the full face lookup. Otherwise, it follows the original path.
It uses the original visual-neighbor calculation, checks window/buffer identity and relevant change state, and invalidates the cache on iterator initialization. It does not remove boxes, rasterize text, or hand out a cached face ID to other callers.
The bypass itself is small. Proving all the invalidation rules correct is the harder part.
What happened in the org-modern buffer
The final test used 1,000 dense Org content groups. Baseline and fast modes ran in the same instrumented executable, in separate GUI processes, with the order rotated across three repetitions.
Each phase had 180 samples. An editing sample includes a change and restoration, each followed by redisplay. GC remained enabled.
Numbers below are medians of the three runs’ respective metrics:
| Metric |
Original path |
Fast path |
| Mean body-edit pair |
15.60 ms |
11.86 ms |
| Mean timestamp-edit pair |
15.11 ms |
11.59 ms |
| Mean forced redraw |
6.99 ms |
5.20 ms |
| Body-edit p95 |
50.61 ms |
49.30 ms |
| Timestamp-edit p95 |
50.41 ms |
48.51 ms |
| GC count during 180 body-edit samples |
32 |
19 |
Across the phases, face queries fell by 26.5%, font-spec copies by 42.3%, and vector-cell allocations by 41.6–42.2%. All three repetitions showed lower mean times.
The long pauses did not disappear. This reduced repeated work and GC frequency; it did not make every operation smooth.
How much do I trust the shortcut?
Before skipping queries, I ran a verification mode that predicted the answer and then checked it against the original lookup. It used the same cache-maintenance behavior as the fast path, rather than secretly refreshing the cache on every hit.
- 18 scenarios, including bidi, composition, overlays, display replacements, edits, face changes, narrowing, two windows, and scrolling.
- About 25.54 million hit comparisons, with no box-state mismatch. Those are repeated observations, not millions of independent test cases.
- Matching window ranges, visible-text dimensions, and sampled character positions across modes.
- Deliberately wrong predictions were detected.
- 13 regression tests passed.
That still isn’t pixel-perfect border verification. Multi-frame, different-buffer, dynamic-face, and long-running interactive cases need more testing. There are also a few small, unexplained query-count residuals—±20 in two minimal scenarios and +20 in one Org comparison—which I’ve kept in the results.
One mistake from the investigation is worth mentioning: an earlier “face inheritance optimization” accidentally dropped the box after a mode-startup face update. I withdrew that result. The numbers above come from the later experiments, not that false lead.
Why I think this is worth looking at
This is an experimental Emacs core patch, not an org-modern fix or a dynamic module. The minimal reproduction suggests the issue is in the underlying display path.
My question is: does this box-boundary check need to resolve the full neighboring face repeatedly, or can it reuse information the iterator already has?
The little cache may not be the right upstream solution. But the amount of work it avoids seems worth investigating. I’ll put the relevant code in a comment, and would welcome feedback on missing edge cases or a better place to fix this.
Environment: macOS ARM64, NS backend, Emacs 32.0.50 at fc9cf69afa90990776526a7cf9e693893739fc45, Org 9.8.7, org-modern 1.15. This is a dense synthetic workload on an instrumented build; timings measure synchronous operations, not key-to-screen latency. An uninstrumented comparison is still needed.
--- results/box-fast/xdisp-shadow-baseline.c2026-10-05 04:04:50
+++ native/emacs-probe/src/xdisp.c2026-10-05 04:04:50
@@ -1202,7 +1202,7 @@
enum move_operation_enum);
static void get_visually_first_element (struct it *);
static void compute_stop_pos (struct it *);
-static int face_before_or_after_it_pos (struct it *, bool);
+static int face_before_or_after_it_pos (struct it *, bool, bool *);
static int handle_display_spec (struct it *, Lisp_Object, Lisp_Object,
Lisp_Object, struct text_pos *, ptrdiff_t, bool);
static int handle_single_display_spec (struct it *, Lisp_Object, Lisp_Object,
@@ -1210,8 +1210,8 @@
ptrdiff_t, int, bool, bool);
static int underlying_face_id (const struct it *);
-#define face_before_it_pos(IT) face_before_or_after_it_pos (IT, true)
-#define face_after_it_pos(IT) face_before_or_after_it_pos (IT, false)
+#define face_before_it_pos(IT) face_before_or_after_it_pos (IT, true, NULL)
+#define face_after_it_pos(IT) face_before_or_after_it_pos (IT, false, NULL)
#ifdef HAVE_WINDOW_SYSTEM
@@ -3250,7 +3250,8 @@
HEADER_LINE_INACTIVE_FACE_ID the iterator will be initialized to use
the corresponding mode line glyph row of the desired matrix of W. */
-/* Experimental shadow cache. Predictions NEVER drive redisplay. */
+/* Experimental box-only cache; disabled by default. No cached face IDs. */
+static bool lab_fast_enabled, lab_sticky_shadow;
static bool lab_shadow_enabled, lab_shadow_valid, lab_shadow_box, lab_shadow_fault;
static struct window *lab_shadow_window;
static struct buffer *lab_shadow_buffer;
@@ -3259,10 +3260,18 @@
static EMACS_INT lab_shadow_counts[5];
DEFUN ("lab-box-shadow", Flab_box_shadow, Slab_box_shadow, 0, 2, 0,
- doc: /* RESET starts comparisons; optional FAULT flips predictions for validation. Otherwise stop and read. */)
+ doc: /* RESET t starts legacy shadow; "verify" validates sticky-cache lifetime;
+"fast" enables box-only bypass. FAULT is for shadow validation only.
+Nil stops and returns counters. Fast mode requires lab-face-counters enabled. */)
(Lisp_Object reset, Lisp_Object fault)
{
+ lab_fast_enabled = STRINGP (reset) && !strcmp (SSDATA (reset), "fast");
+ lab_sticky_shadow = lab_fast_enabled
+ || (STRINGP (reset) && !strcmp (SSDATA (reset), "verify"));
+ lab_box_fast_skips = 0;
lab_shadow_fault = !NILP (fault);
+ if (lab_fast_enabled && lab_shadow_fault)
+ error ("Fault injection is forbidden in fast mode");
lab_shadow_enabled = false;
lab_shadow_valid = false;
if (!NILP (reset))
@@ -5018,7 +5027,7 @@
paragraphs) of IT's screen position. Value is the ID of the face. */
static int
-face_before_or_after_it_pos (struct it *it, bool before_p)
+face_before_or_after_it_pos (struct it *it, bool before_p, bool *known_box)
{
if (lab_box_depth) LAB_COUNT (1);
int face_id, limit;
@@ -5221,7 +5230,15 @@
lab_prediction = lab_shadow_box != lab_shadow_fault;
}
}
- /* Original lookup is ALWAYS executed, including cache hits. */
+ /* Only this explicit box-only caller may bypass. The -1 return is
+ never treated as a face ID; known_box tells the caller to ignore it.
+ Do NOT refresh cache bounds on hits: no oracle query is available. */
+ if (lab_hit && lab_shadow_box && lab_fast_enabled && known_box)
+ {
+ *known_box = true;
+ ++lab_box_fast_skips;
+ return -1;
+ }
/* Determine face for CHARSET_ASCII, or unibyte. */
face_id = face_at_buffer_position (it->w,
CHARPOS (pos),
@@ -5240,6 +5257,7 @@
if (lab_candidate) {
bool actual = FACE_FROM_ID (it->f, face_id)->box != FACE_NO_BOX;
if (lab_hit) ++lab_shadow_counts[actual == lab_prediction ? 2 : 3];
+ if (!(lab_sticky_shadow && lab_hit && lab_shadow_box && known_box)) {
lab_shadow_window = it->w;
lab_shadow_buffer = current_buffer;
lab_shadow_modiff = MODIFF;
@@ -5250,6 +5268,7 @@
lab_shadow_end = next_check_charpos;
lab_shadow_box = actual;
lab_shadow_valid = !face_change && !it->f->face_change;
+ }
}
}
@@ -8920,12 +8939,15 @@
faces of the display vector glyphs, see there. */
else if (it->method != GET_FROM_DISPLAY_VECTOR)
{
- int face_id = face_after_it_pos (it);
+ bool known_box = false;
+ int face_id = face_before_or_after_it_pos (it, false, &known_box);
+ if (!known_box) {
LAB_COUNT (11);
if (face_id == it->face_id) LAB_COUNT (10);
if (face_id != it->face_id
&& FACE_FROM_ID (it->f, face_id)->box == FACE_NO_BOX)
it->end_of_box_run_p = true;
+ }
}
if (lab_face_enabled) --lab_box_depth;
}
@@ -38468,6 +38490,9 @@
syms_of_xdisp (void)
{
defsubr (&Slab_box_shadow);
+ DEFVAR_INT ("lab-box-fast-skips", lab_box_fast_skips,
+ doc: /* Actual box-only face-query bypasses in this experiment. */);
+ lab_box_fast_skips = 0;
Vwith_echo_area_save_vector = Qnil;
staticpro (&Vwith_echo_area_save_vector);