跳转到主要内容

性能

本节将讨论嵌入式图形用户界面的性能。

这里的高性能是指在获得所需图形效果和动画时实现高帧率。

我们回顾一下上一节中关于主循环如何影响用户界面帧率的内容。再次假设有一台连接到LTDC的并行RGB显示屏和两个帧缓冲。基本情况如下:

双帧缓冲

假设显示屏每秒刷新60次,则每次刷新之间的间隔约16 ms。 The calculation is this:

1 s / 60 = 0.01667 s = 16.67 ms.

在1号帧缓冲的传输开始时,TouchGFX开始将1号帧绘制到2号帧缓冲。如果1号帧的渲染在下一次传输开始前完成,则可以传输2号帧缓冲。如果没有在16.67 ms内完成,则再次传输1号帧缓冲,并且显示屏的显示内容不变:

主循环时间超过16.67 ms

这种情况意味着丢帧。

The time for the collect and update phases are typically minuscule, e.g. less than 1 ms, and therefore more or less negligible when considering the overall time taken of the main loop. 因此,在下文和全文中,在考虑渲染时间时,其中包括采集和更新阶段。

如果许多帧的渲染时间超过16.67 ms的时限,显示屏上的帧率将是30帧每秒(fps)。

如果渲染时间大体上短于16.67 ms,但有一些帧超过16.67 ms,则平均帧率可能接近60 fps,但在用户看来动画可能不流畅。动画中的某些步骤可能看起来过快而某些又过慢,具体取决于应用。 This is not desirable.

渲染时间还可能更长。 If it is just above 33 ms, the frame rate will drop to 20 fps as we only have a new frame ready on every third transfer.

FPS(帧率)最长渲染时间
6016.67 ms
3033.34 ms
2050.00 ms
1566.67 ms

The table shows the maximum rendering time (including the collect and update phases) that is available for a given frame rate.

为了获得良好的用户界面性能,最好定期检查和监测帧率。可使用两种方法:

  • 测量渲染时间
  • 丢帧计数

测量渲染时间​

测量渲染时间的第一种方法提供了最详细的信息。其理念实际上是测量从帧传输到渲染阶段结束之间的时间。图形引擎在采集阶段开始时调用GPIO类的函数,并在渲染阶段结束时再次调用。 The application defines these functions and can hook into them to perform measurements.

测量可通过两种方式来执行:

  • Use an external timing device like an oscilloscope: To measure using an oscilloscope, the application should implement the set(GPIO_ID) and clear(GPIO_ID) methods from the GPIO interface. 然后,示波器可以测量输出为高电平的持续时间,以此作为渲染时间。
  • 使用内部计时器:另一种方法是使用内部计时器,如sysTick计时器。在调用GPIO::set(RENDER_TIME),应用可将计时器的值保存在变量中。在进行清零调用时,应用可再次读取计时器并减去前值,从而获得渲染时间。计时器的速度将决定测量精度。应用必须以某种方式使渲染时间可见。一种方式是将值保存在全局变量中,并且可以在界面上的TextArea中显示值。也可使用调试器检查值。

丢帧计数​

图形引擎对上一个采集-更新-渲染阶段中发生的传输次数进行计数。应用可轻松地检查此值,以便了解是否丢帧以及帧率是否下降。

计数在HAL类中提供:

void handleTickEvent() {
tickCounter += 1;
if (HAL::getInstance()->getLCDRefreshCount() > 1) {
//Alert programmer somehow
...
}
}

丢帧补偿​

When frames are lost and the frame rate of one of our animations therefore lowered we can compensate to a certain degree. 我们可以:

  • 等待其结束 - 让动画继续,这会导致动画持续时间变长,并且动画可能不流畅。
  • skip some frames - make sure that the overall animation does not take more time than intended by skipping frames.

TouchGFX can be instructed to automatically skip some frames, when frames are lost. 可通过在每个实际帧将动画tick一次以上来实现。当渲染时间不稳定时,这有助于让动画更流畅。

HAL.hpp
void setFrameRateCompensation(bool enabled)

影响渲染时间的因素有哪些?​

影响渲染时间的因素有许多:更新部分的大小、分层的使用、控件的复杂度和渲染可使用的硬件支持。

界面更新了多少?​

渲染时间通常与必须更新的像素数成正比。如果动画需要的渲染时间过长,一种可能的解决办法是缩小动画面积。例如,如果有一张旋转图像而性能不够好,则可通过缩小图像尺寸来改善性能。

通过缩小图像尺寸缩短渲染时间

记住,图形引擎会重绘应用使之无效的区域。 This means that it is important to only invalidate the areas that actually require a refresh.

无效区域越大,渲染时间越长。

图形中的层数​

在典型应用中,图形将包含彼此堆叠的不同元素。如果更新了元素中的一个,通常必须重绘所有元素。

典型的例子是背景图像、帧和一些文本:

分层图形元素

此用户界面的创建方法是将TextArea控件放在显示了一个透明帧的Image控件上方。二者都在背景图像上层:

TouchGFX Designer中的分层图形元素

This solution is used very often in applications. 这是一种十分简单的解决方案,具有高度的灵活性,例如,可以在运行时间更改帧或在背景上移动帧和文本。

关于渲染时间的问题是如果在运行时间更新了文本且需要重绘,图形引擎还需要重绘背景和帧,然后是新的文本。这会显著增加文本的渲染时间。

无效区域的层数越多,渲染时间就越长。

渲染像素的复杂度​

将每个像素渲染到帧缓冲的难度并不一致。在所有类型的渲染中,图形引擎必须将最终的像素写入帧缓冲。但是,要写入像素的计算需要的消耗并不相同。

固定色彩(如Box Widget中使用的色彩)的消耗最低,只需计算一个像素并将结果重复用于所有像素。这意味着使用许多Box可获得非常高的性能。由于这会导致用户界面质量不高,因此不建议这样做。

第二低的是图像的像素计算消耗,这是因为像素均以可直接使用的格式存储在位图中。计算要写入帧缓冲的像素关系到从位图中的正确位置加载色彩值。

文本的消耗与图像相当,每个字母实际上都是一幅小图像。事实上,由于大量小图像导致了相当高的“开始-停止”消耗,因此文本的消耗更高。例如,每个字母的位置计算。为了让文本看起来尽可能美观,会将文本显示为具有透明度的小图像,请参见下文关于透明的注释。

旋转或缩放后的图像消耗更高。任务同样是从位图加载像素值,但由于图形引擎必须包含缩放和旋转,因此这时的计算更耗时。

几何元素(如圆)的消耗比之更高。这时我们不能从位图加载像素色彩,而是必须计算圆的形状和圆中每个像素的色彩。

透明度增加了元素的绘制消耗。如果一些像素不是实心的,那么元素是透明的。图形引擎首先必须绘制透明元素后方的元素(如“帧中的文本”部分所述),这会增加绘制的消耗。其次,图形引擎随后必须将背景像素与透明元素的像素进行组合,并将结果写入帧缓冲。此类计算的耗时显著多于只写入计算过的像素的场景。

Box、Image、旋转Image和圆。实心元素位于第一行。透明元素在下方。

透明总是需要多出一层。但是,将实心像素放在其他实心像素的上方并不一定会增加层数。 The graphical engine tries not to draw pixels that are covered by other solid pixels, as this would be a waste of precious time.

无效区域中元素的复杂度越高,渲染时间就越长。

Remember that it is only the elements that are part of the invalidated area that add to the rendering time. 无效区域之外的元素对渲染时间无影响。

点击此处阅读关于UI组件和性能的更多内容。

渲染的硬件支持​

Some STM32 microcontrollers contain graphical accelerators. These accelerators can reduce rendering time, as the accelerators can run in parallel with the microcontroller core. The core will then be able to run other tasks, while the accelerators render graphics.

The accelerators are automatically used by TouchGFX when available.

何时应考虑渲染时间​

渲染时间并非总是那么重要。当低帧率可被用户观察到时,应注意渲染时间。当动画在屏幕的一部分上运行(如旋转的图标)或您在界面上移动或滑动某元素时,通常就属于这种情况。如果更新频率低,那么在用户看来,动画将呈现出分步显示而非流畅的状态。如果是这样,应检查渲染时间。

另一方面,如果用新界面替代整个界面,当更换期间帧率显著下降时,用户通常注意不到。这是因为用户看不到渲染何时开始,只能看到它何时结束。

这两条规则意味着对于动画元素(如移动元素)而言,应使用较少的层数,避免使用复杂元素和许多层数。对于界面的其余部分,这些不是问题。

模拟时钟和滚动列表

在本例中,左侧有一个模拟时钟。通过旋转三幅细长的图像渲染三根指针。这通常不难实现,因为指针并非总是在移动。但如果我们要让时钟在界面上到处移动,将会在每一帧中重绘指针,由于绘制旋转图像通常比较耗时,因此会比较复杂。

右侧是一个滚动列表。 The user can move this list of numbers up and down, so we need a high frame rate for the user interface to appear responsive. 因此,必须考虑滚动列表中元素的渲染时间,或者缩小滚动列表的尺寸。

通过使内容无效来优化性能​

通常整个控件均无效,但图形引擎只能使控件的内容无效,而不能使整个控件无效。通过减少无效区域,渲染时间一般会明显缩短。渲染时间的改进取决于:

  • 控件内容覆盖的区域相对于整个控件的大小。
  • 背景控件部分或全部被控件覆盖。

The following figures illustrate the concept of invalidating content, by using the TextArea widget as an example. 图1显示了控件的整个区域。图2显示了使用TextArea::invalidate()时的无效区域。图3显示了使用TextArea::invalidate()时的无效区域。

图1. 横跨整个屏幕宽度的文本区域

图2. 使用TextArea::invalidate()时无效的区域(红色)

图3. 使用TextArea::invalidateContent()时无效的区域(绿色)

使用TextArea::invalidateContent()的示例​

在控件与其他控件重叠的情况下,使用TextArea::invalidate()使整个TextArea无效时,需要重新绘制其他控件。通过改用TextArea::invalidateContent(),我们将不必要的无效和重绘控件的风险降至最低。这对于昂贵的控件尤其如此,例如:圆、仪表等。

下图说明了如何避免使用TextArea::invalidateContent()使背景控件(图像-意法半导体标志)无效。如果我们使用TextArea::invalidate(),将会使后台控件无效并重新绘制。

使用TextArea::invalidateContent()的示例

获得良好性能的建议​

我们总结了获得良好性能的建议,以结束本节内容:

  • 不要重绘未更改的部分 确保没有误操作将界面上不必要的部分失效。这会降低性能且无任何益处。
  • 在质量与速度之间寻求平衡 降低元素的复杂度有助于提高性能。复杂度与性能之间的良好平衡通常极为关键。
  • Utilize hardware capabilities The capability of a microcontroller with hardware acceleration is often higher than a microcontroller without. Consider using a microcontroller with hardware acceleration.
  • 用图像替代计算图形 计算得到的圆比圆图像慢。一般而言,图像可替代许多静态元素。
  • 调整显示屏刷新率 如本节开头所述,刷新率是渲染时间的硬性限制。如果渲染时间超过刷新率,帧率将下降。如果渲染时间只超过刷新率一点点,也许能够将显示屏的刷新率降至如55 Hz(相当于18.2 ms)这样的水平,并维持高帧率。

Framebuffer strategies​

The Partial Framebuffer, Emulated Framebuffer, and Hybrid Double Buffer strategies render updated graphics in small slices. To keep rendering time low, reduce the use of widgets that are expensive to draw in multiple small areas, such as TextArea with automatic line wrapping, Circle, Video, and SVGs.
Widgets such as Box, image based Widgets, and TextureMapper have very little overhead in this rendering pattern.

For the Emulated framebuffer strategy it is important to have an even load. It is adviced to separate expensive elements.

Graphics Accelerators​

A common way to achieve high performance is to use an STM32 MCU with a graphics accelerator. The STM32 MCU family mainly uses two graphics accelerators: Chrom-ART and NeoChrom GPU. Chrom-ART is also called DMA2D, and NeoChrom GPU is also called GPU2D.

The Chrom-ART accelerator is capable of copying and blending images between the image storage and the frame buffer. NeoChrom GPU can do the same, but is also capable of texture mapping and vector graphics.

Chrom-ART comes in three versions. The first version of Chrom-ART is capable of copying and blending. The second version is also capable of converting YCbCr data to RGB. This improves the JPEG decoding performance. The third version is called Chrom-ART2 (DMA2D3) and is command-list based, supports more image formats, and is capable of down-scaling with nearest-neighbor.

Compared to Chrom-ART, NeoChrom GPU is capable of accelerating more graphical operations and has a richer feature set.

NeoChrom GPU comes in two versions: NeoChrom GPU and NeoChromVG GPU. NeoChromVG GPU has the same feature set as NeoChrom GPU, but also has full hardware acceleration of vector rendering.

Graphic featureChrom-ARTChrom-ART2NeoChrom GPUNeoChromVG GPU
Supported formats (with TouchGFX)ARGB8888, RGB888, RGB565, A8, A4, L8-888, L8-8888ARGB8888, RGB888, RGB565, A8, A4, A2, A1, L8-565, L8-888, L8-8888ARGB8888, RGB888, RGB565, A8, A4, A2, A1ARGB8888, RGB888, RGB565, A8, A4, A2, A1
Command list basedNoYesYesYes
DrawingRectanglesRectanglesRectangles, Pixels, Line, Triangle, Quadrilaterals with multi-sample anti-aliasing (MSAA)Rectangles, Pixels, Line, Triangle, Quadrilaterals with multi-sample anti-aliasing (MSAA)
BlittingCopy, alpha blending, pixel format conversionCopy, alpha blending, pixel format conversionCopy, alpha blending, pixel format conversion, color keyingCopy, alpha blending, pixel format conversion, color keying
ScalingNoYes (down-scaling with nearest neighbor)YesYes
Texture MappingNoNoYesYes
Vector GraphicsNoNoNo*Yes

* Vector Graphics are partially hardware-accelerated with NeoChrom GPU when using TouchGFX. Full hardware acceleration for Vector Graphics is available with NeoChromVG GPU.

With these capabilities available many TouchGFX Widgets are accelerated:

WidgetChrom-ARTChrom-ART2NeoChrom GPUNeoChromVG GPU
Box, BoxWithBorderYesYesYesYes
Image, AnimatedImage, TiledImage, SnapshotWidgetYesYesYesYes
Button, ButtonWithIcon, ButtonWithLabel, ToggleButtonYesYesYesYes
RadioButton, RepeatButtonYesYesYesYes
PixelDataWidgetYesYesYesYes
TextArea, TextAreaWithWildcard, KeyboardPartlyYesYesYes
ScalableImageNoPartlyYesYes
TextureMapper, AnimatedTextureMapperNoNoYesYes
Circle, Line, GraphNo*No*No*No*
SVGNo*No*No**Yes

* The drawing/blending of pixels to the framebuffer is done by Chrom-ART or NeoChrom GPU, but the shape calculations are done in software.
** SVG rendering is partially hardware accelerated with NeoChrom GPU and fully accelerated with NeoChromVG GPU.

The operations that are not hardware-accelerated are software-rendered (implying a higher CPU load, and lower performance). As the table above shows, NeoChrom GPU can accelerate widgets like ScalableImage and TextureMapper. This means that we can use those widgets to a greater extent while keeping a high performance.

Read more about TouchGFX with NeoChrom here.