InspireFaceInspireFace1.2.4.d3
Home
Get started
Get and build the SDK
Examples
  • English
  • 简体中文
GitHub
Home
Get started
Get and build the SDK
Examples
  • English
  • 简体中文
GitHub
  • Introduction
  • Get started
  • Features
  • Guides

    • Architecture and lifetime
    • Model packs
    • Image inputs and coordinates
    • Sessions and tracking
    • Face analysis
    • Recognition and FeatureHub
    • Facial landmarks
    • Liveness detection
    • Face capture
    • More API recipes
  • Language and platform

    • C API
    • C++
    • Python
    • Android
    • Apple
    • iOS
    • macOS
    • HarmonyOS
  • Get and build the SDK

    • Overview and downloads
    • Source and common options
    • Linux
    • macOS
    • Android
    • iOS
    • HarmonyOS
    • NVIDIA TensorRT
    • Rockchip NPU
    • Python packaging
  • Hardware deployment

    • ARM
    • NVIDIA TensorRT
    • Rockchip NPU
    • Python on Rockchip
  • InspireCV
  • Complete examples
  • API coverage
  • Performance
  • Image processing benchmarks
  • Troubleshooting

ARM

InspireFace runs on ARM CPUs in mobile devices, embedded Linux systems and Apple Silicon computers. ARM optimization covers image preparation, model inference and feature comparison, while session reuse keeps repeated work out of the frame loop.

The SDK uses InspireCV for image operations such as resizing and channel conversion, with ARM NEON paths for supported formats. The CPU inference backend is also adapted to ARM instructions and compute kernels to accelerate model execution. The SDK reduces additional work in input preparation and feature comparison.

1 · Camera inputPixel format, stride and orientation
2 · PreprocessingResize, align and prepare model inputs
3 · CPU inferenceARM-adapted model execution
4 · ResultsTracking, features and similarity

Image and feature operations on ARM

The default InspireCV Image backend provides vectorized paths for supported data types and channel counts. NEON processes several values per instruction; the kernels also reduce repeated coordinate calculations and organize pixel reads and writes for common transforms.

OperationImplementationWhere it is useful
ResizeReuse precomputed horizontal sample indices and interpolation weights; use NEON in supported kernels.Detection previews and fixed-size model inputs.
WarpAffineSeparate identity, scale/translation and general affine paths; vectorize supported sampling operations.Face alignment and transformed image regions.
Rotate / flipUse block transposes and vector loads/stores for supported formats.Camera orientation and preview preparation.
Color conversionVectorized RGB/BGR exchange and grayscale conversion.Matching the channel order needed by the next stage.
Feature comparisonThe SDK's ARM NEON dot product processes four float components at a time.Similarity between face embeddings.

The image operators retain general CPU implementations for other combinations and scalar handling for the remainder of a vector block. The selected path depends on the operation, pixel type, channels and build target. The InspireCV guide shows the image APIs and their coordinate rules.

Reuse work across frames

Several SDK optimizations help ARM deployments as well as other CPU platforms:

  • Skip unnecessary resizing. When an input already matches the model's required dimensions, the adapter uses a pixel view and skips the resize step.
  • Reuse conversion and tensor objects. The inference adapter caches input/output host tensors and image converters, rebuilding them when the relevant shape or configuration changes.
  • Read detection outputs through views. Detection postprocessing reads the inference output directly through borrowed views, avoiding an extra copy into intermediate vectors.
  • Keep a tracking session alive. LIGHT_TRACK reuses previous-frame information and runs detection when needed. See tracking modes and latency for the tradeoff between frame cost and discovering new faces.

Keep the session and model alive across a video sequence. Borrowed pixels and result views still need valid storage until their consumers finish; ownership and lifetime explains when to retain or copy them.

Optional Task preprocessing

InspireCV Task brings geometric sampling, color conversion, normalization and tensor layout into one preprocessing flow. Its general execution path works in small tiles with reusable temporary buffers, reducing the need to create a full intermediate image for every stage.

StageAvailable work
Camera formatsNV12, NV21 and I420 sampling and conversion to RGB-family outputs.
Tensor valuesConvert bytes to float and apply per-channel (value - mean) × scale.
Tensor layoutWrite HWC or CHW output directly into the caller's tensor buffer.
Repeated executionReuse pipeline configuration and existing output storage with RunInto or TensorBuffer.

These are Task capabilities for application preprocessing. InspireFace can also use Task for its camera-stream preprocessing when built with ISF_ENABLE_INSPIRECV_TASK_PREPROCESS=ON; that SDK option is off by default. The switch selects the stream preprocessing backend. Model-specific normalization and tensor preparation still follow the model adapter. See Task examples to use the API directly.

Build the Task path on ARM64 Linux

After preparing the source, run this on an ARM64 Linux device with a native compiler. It builds a shared CPU SDK:

cmake -S . -B build/arm-cpu-task \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_POLICY_VERSION_MINIMUM=3.5 \
  -DISF_BUILD_SHARED_LIBS=ON \
  -DISF_BUILD_WITH_SAMPLE=OFF \
  -DISF_BUILD_WITH_TEST=OFF \
  -DISF_ENABLE_INSPIRECV_TASK_PREPROCESS=ON \
  -DINSPIRECV_TASK_ENABLE_ARM_NEON=ON
cmake --build build/arm-cpu-task --parallel 4
cmake --install build/arm-cpu-task

The SDK is installed at build/arm-cpu-task/install/InspireFace. For cross-compilation or mobile targets, use the corresponding platform build guide and apply these options to that toolchain configuration.

INSPIRECV_TASK_ENABLE_ARM_NEON defaults to ON and controls the explicit Task NEON paths. It does not control all Image operators or compiler auto-vectorization. NEON support is selected at build time; ARMv7 binaries compiled with NEON require a processor that supports those instructions. Compare the Task and default preprocessing paths using the same input formats and transformations on your device.

Choose the SDK for the device

Use a CPU resource pack such as Pikachu or Megatron with the CPU SDK. Download links are collected in SDK overview and downloads. Select the OS, process architecture and C runtime together:

TargetArchitecture / ABIBuild and integration
LinuxARM64 aarch64; ARMv7 hard-floatBuild · C API · Python packaging
Androidarm64-v8a; armeabi-v7aBuild · Camera integration
iOSDevice arm64; simulator arm64 on Apple SiliconBuild · Objective-C / Swift integration
HarmonyOSarm64-v8aBuild · ArkTS integration
macOS Apple SiliconNative arm64 processBuild · macOS · C++ · Python

For Linux, match glibc/uClibc and the compiler runtime to the device image. On Apple Silicon, an x86_64 process running through Rosetta needs an x86_64 SDK. Rockchip NPU deployment has its own SDK, model and driver requirements.

Keep the camera loop efficient

  1. Match the actual input layout. Pass the camera's supported pixel format instead of repeatedly converting between YUV, RGB and BGR. The raw C stream requires tightly packed rows; repack padded or plane-based input as described in image inputs.
  2. Limit the work to the application. Choose a supported detector size, adjust the preview size, set a realistic face limit, and enable only the analysis outputs you use. Check small-face recall when reducing dimensions.
  3. Preserve one sequence per session. Reuse the session and keep frames in order. A short queue prevents old camera frames from adding display latency.
  4. Measure a sustained run. Record preprocessing, tracking, optional analysis and total frame time after warm-up. Compare median and p95, then check performance after the device has warmed up thermally.

NEON speedups depend on the operation, compiler and CPU. Compare both output correctness and timing when changing preprocessing paths. The performance guide provides complete timing examples; the image-processing benchmarks explain how to separate image operations from complete SDK calls.

Edit this page
Last Updated:: 9/28/26, 8:17 PM
Contributors: Jingyu
Next
NVIDIA TensorRT