Artificial Int News
2026-10-12

Daily AI News - October-12-2026

From 182 items, 5 important content pieces were selected

Microsoft Releases AesCode 8B and 32B for Visual HTML/CSS Generation ⭐️ 7.48/10

Microsoft has released AesCode in 8B and 32B sizes, models that generate editable HTML/CSS artifacts such as slides, posters, and dashboards. AesCode uses an image generated from the same prompt as a visual reference while following the prompt’s requirements for content. The models aim to combine the visual composition strengths of image generation with the structured, editable output of code generation. This could help people create polished visual materials while retaining control over their text, numbers, and layout. AesCode separates semantic requirements from visual cues using graph-structured supervision and decoupled cross-modal rewards. AesCode-32B is based on Qwen3-VL-32B-Instruct, is trained with cold-start SFT followed by GDPO across seven reward channels, and leads the reported aggregate benchmark results.

reddit · r/LocalLLaMA · /u/jacek2023 · Oct 11, 11:17

Background: HTML and CSS are web technologies used to describe a page’s structure and visual styling, so their output can be edited as code rather than being only a flat image. The challenge described here is that code models may not understand how layout, hierarchy, and color appear on a canvas, while image generators can struggle to render exact text, numbers, and logical relationships.

Tags: #high value

OpenMed 3.0 Adds Local Patient Timelines ⭐️ 7.4/10

OpenMed 3.0, an Apache-2.0 medical AI toolkit, adds Patient Journey, which combines clinical notes, FHIR and HL7 feeds, lab files, and imaging reports into a patient timeline on local hardware. The project says the release received 148 pull requests, including 46 community contributions, and OpenMed 3.1 is under development with 422 open issues. The release targets a practical healthcare problem: patient information is scattered across systems and sources can conflict, while the toolkit is designed to keep data and inference on the user's own hardware. This could help developers build privacy-conscious clinical workflows, though the project explicitly says its 3.0 summarization pipeline is synthetic-only and not ready for real patients. OpenMed says it refuses to auto-merge even a 99% identity match, preserves conflicting diagnoses and corrected lab values for review, and will not silently call a cloud model if a requested local model is unavailable. Its demo uses a synthetic patient and combines five synthetic sources in about 0.07 seconds; the release adds no new clinical model checkpoints.

reddit · r/LocalLLaMA · /u/dark-night-rises · Oct 11, 13:16

Background: FHIR is a standard for exchanging healthcare information electronically, while HL7 is the standards organization behind FHIR and also refers to healthcare data-exchange formats. OpenMed describes support for model use through options including ONNX, GGUF with llama.cpp, and browser-based WebGPU or WebAssembly, rather than requiring a hosted AI service.

References

Tags: #high value

Project Maya Brings GLM-5.3-Flash to Local GPUs ⭐️ 7.38/10

A Reddit user reports that Project Maya’s Strata-based setup increased their GLM-5.3-Flash generation speed from about 10 to 30 tokens per second. The project is designed to run the 321-billion-parameter model on local hardware by distributing its workload across GPU memory, system RAM, and NVMe storage. If the reported performance is reproducible, Maya could make a very large model practical for more people without requiring datacenter-scale GPU memory. Its local operation and OpenAI- and Anthropic-compatible APIs may also appeal to users seeking private inference and compatibility with existing tools. GLM-5.3-Flash has 321 billion total parameters, with about 18 billion active per token; the project says its compact Maya-S quant is 96.5 GB and reports 97.7–99.2% of the FP8 model’s zero-shot accuracy. The Reddit post’s 30-token-per-second result is an individual report, while the project lists AMD and Windows support as experimental.

reddit · r/LocalLLaMA · /u/inthesearchof · Oct 11, 18:42

Background: GLM-5.3-Flash is a mixture-of-experts model: it contains many specialist components, but only a subset is active for each token. Expert offloading reduces the amount of model data that must fit in GPU memory by keeping frequently used weights on the GPU and placing others in system RAM or NVMe storage, though those slower tiers can affect speed.

References

Discussion: The poster is enthusiastic, saying the speed improvement makes GLM-5.3-Flash usable and that they currently prefer its feel to Qwen 3.8 Flash Next, even at a low quantization level. They note that testing is still ongoing, and the post does not provide hardware details for the reported 30-token-per-second result.

Tags: #high value

Qwen 3.8 27B Megakernel Reaches 140 Tokens per Second on an RTX 3090 ⭐️ 7.15/10

The OpenJet CUDA megakernel’s follow-up benchmarks report 140 tokens per second for code generation on a single RTX 3090, compared with 73 tokens per second for llama.cpp with speculative decoding. At 1K context, its mean KL divergence against the reference was 0.0009, versus 0.0019 for llama.cpp. The results suggest that a GPU-specific inference engine can substantially increase local Qwen 3.8 27B throughput while closely matching llama.cpp’s output distribution in the reported tests. This may benefit developers running large models on consumer GPUs, though the results are tuned to one RTX 3090 and are not yet backed by task-level pass-rate benchmarks. The engine reports 78 versus 51 tokens per second in reasoning mode, 73 versus 52 with a 30K-token code context, and roughly 1,600 versus 1,100 tokens per second for prompt processing. It is tuned on an RTX 3090, supports a limited set of Q4_K_M model files natively or semi-compatibly, and falls back to llama.cpp for other quantizations; speculative decoding is the main source of the speedup.

reddit · r/LocalLLaMA · /u/Adorable_Weakness_39 · Oct 11, 13:04

Background: A megakernel combines multiple stages of GPU computation into a single kernel, which can reduce the overhead of launching many separate GPU operations. LLM inference can benefit from this approach when repeated launch overhead is significant, although actual speed depends on the model, hardware, and implementation.

References

Tags: #high value

Xiaomi Details MiMo-V2.6’s Costly, Large-Scale RL Training ⭐️ 7.08/10

Xiaomi has disclosed a large-scale reinforcement-learning (RL) post-training approach for MiMo-V2.6, with the headline reporting a cost of $2.6 million for one run and 3.7 billion tokens per step. Search results also describe public training runs for the MiMo-V2.6-Pro and MiMo-V2.6-Flash models. The reported cost and token volume illustrate the substantial compute and data throughput involved in scaling RL for large language models. Xiaomi’s public disclosure could give researchers and developers useful visibility into training practices, although the figures alone do not establish model quality or efficiency. The headline gives a per-step volume of 3.7 billion tokens, while a separate search result describes a public training event as costing $3.47 million over six days and involving 7,000 task environments; these may reflect different scopes or accounting, and the available results do not reconcile them. Xiaomi’s live RL page provides training metrics from trainer logs, but the search results say final specifications, APIs, and pricing had not yet been announced.

rss · InfoQ 中文站 · Oct 10, 17:24

Background: RL post-training is a stage in which a model is further trained using reinforcement-learning methods after its initial training. A token is a unit of text processed by a language model, so tokens per step indicate the volume handled during each training step. Xiaomi’s MiMo-V2.6 RL page reports live metrics for the Pro and Flash training runs.

References

Tags: #high value

Previous Briefings