GLACIER AI · PRODUCT INFO

AI Info Center

Berserk is Glacier AI's core in-house model series. AI Tools provides task tools, and the Aggregation Hub provides model access.

The in-house Berserk model series

Includes chat, text-to-image, text-to-video and other single-modal and cross-modal models.

Positioning: results close to the benchmark model, with a token price of 50% of the corresponding benchmark's lowest market price of the day. Reference prices are updated or may change daily; the actual price is the day's quote. This describes the pricing rule, not a live market quote.

Aggregation Hub

Unified access to Chinese models, with one protocol and keys and calls enabled per project.

Integrations cover Qwen, ERNIE, Zhipu, Kimi, DeepSeek, Doubao, Hunyuan and Spark; specific model versions and APIs are confirmed at integration.

AI Tools

84 features in six categories, covering video, text, images, speech, cross-modal, time series and perception. Super-resolution, enhancement, AI denoising and bitrate recovery are part of AI Tools.

View original image and video demos ↗

Full feature catalog

Text-to-video and image-to-video

Generate video clips from text prompts or reference images for creative content, ad assets and previsualization.

Video super-resolution and frame interpolation

Upscale video resolution and frame rate for smoother, sharper footage, ideal for restoring old films and slow motion.

Video highlight detection and auto-editing

Automatically find the best moments and cut them into short clips for sports, live streams and short-form video.

Video subtitle removal

Detect and remove burned-in subtitles for re-dubbing and multilingual distribution.

Video watermark removal

Locate and remove watermarks and channel logos for asset cleanup and compliance.

Video stabilization

Remove handheld camera shake for steady footage and a better viewing experience.

Video matting and background replacement

Separate people in the foreground from the background, enabling virtual backgrounds without a green screen.

Action and event recognition

Recognize human actions and events in video for security, sports and industrial monitoring.

Activity, fall and fatigue detection

Detect falls, fatigue and abnormal activity in elder care, construction sites and driving, and raise alerts.

Visual anomaly detection

Spot deviations from normal patterns in video for production lines, inspections and security.

Multi-object tracking

Continuously track multiple objects in video and output track IDs for pedestrian and traffic flow analysis.

Object detection and multi-object tracking

Detect objects in video and images and track them over time, powering counting, speed measurement and behavior analysis.

Trajectory prediction and motion planning

Predict the future paths of pedestrians, vehicles and robots for autonomous driving and robot obstacle avoidance.

Lip sync and digital human presenters

Align a digital human's lip movements with speech for virtual hosts, newscasts and customer service.

Vision-language-action and grasping

Drive robotic grasping from vision and language instructions for embodied AI.

General writing and chat

Text generation, rewriting, Q&A and multi-turn chat for general assistant use cases.

Summarization and rewriting

Condense long text, rewrite in a new style and extract key points for reports, news and meeting notes.

Machine translation and localization

Translate between languages and polish for local audiences across documents, interfaces and subtitles.

Text content moderation

Flag violating, violent or extremist, sexual and spam text for communities and content platforms.

Text embeddings and semantic similarity

Encode text as vectors for semantic search, deduplication, clustering and similarity matching.

Entity and relation extraction

Extract people, organizations, events and their relationships from text to build knowledge graphs.

Sentiment and opinion analysis

Determine sentiment polarity and opinion in text for public opinion monitoring, reviews and customer service QA.

Intent recognition and slot filling

Understand user intent and extract key slots for dialogue systems and ticket routing.

Tabular classification and risk scoring

Classify tables and forms and output risk scores for risk control, compliance and review.

Field and form extraction

Extract structured fields from contracts, receipts and forms to automate data entry.

Text language identification

Automatically identify the language of text for routing, pre-translation and content distribution.

Language identification

Identify the language of speech or text for multilingual customer service and meeting systems.

PII/PHI detection and redaction

Detect and redact personal and health information to meet privacy compliance.

AI security detection and response

Identify risks such as prompt injection, jailbreaks and harmful content, and trigger response policies.

Causal inference and uplift modeling

Estimate causal effects and incremental gains of interventions for marketing, operations and pricing.

Graphs, anti-fraud and relationship risk

Mine fraud rings and linked risks from graphs for anti-fraud and risk control.

Text-to-image and reference-based generation

Generate images from text or reference images for creative work, design and marketing assets.

Inpainting, outpainting and controllable editing

Repaint regions, extend beyond the frame and edit with precise control for retouching, e-commerce and design.

Image classification and fine-grained recognition

Recognize image categories and fine-grained subclasses for quality inspection, archiving and identification.

OCR text recognition

Recognize text in images, including receipts, ID documents, documents and scene text.

Layout, table and formula parsing

Parse document layouts, table structures and math formulas for document digitization.

Face detection, recognition and liveness

Face detection, matching and liveness checks for access control, identity verification and security.

Pose, gesture and keypoints

Recognize body pose, hand gestures and keypoints for interaction, fitness and animation.

Medical image detection and segmentation

Detect and segment lesions and organs for diagnostic support and research.

Supervised defect detection

Detect surface and internal product defects from labeled samples for industrial quality inspection.

Visual measurement and alignment

Measure dimensions, position and alignment for precision manufacturing and assembly.

Open-vocabulary detection & segmentation

Detect and segment objects described in natural language, with no fixed classes.

Image & video segmentation

Semantic, instance and panoptic segmentation for matting, labeling and analysis.

3D reconstruction & digital twins

Reconstruct 3D scenes from images or point clouds for digital twins, mapping and simulation.

Depth estimation

Monocular and stereo depth estimation for autonomous driving, AR and robotics.

Super-resolution, denoising & deblurring

Image enhancement and restoration for sharper quality and old photo repair.

Text/image to 3D assets

Generate 3D models and assets from text or images for games, e-commerce and the metaverse.

Speech-to-text ASR

Transcribe speech into text for meetings, customer service and subtitles.

Text-to-speech TTS

Turn text into natural speech for announcements, audio content and accessibility.

Real-time voice agents

Real-time voice conversations that take business actions, for call centers and voice assistants.

Speech-to-speech translation

End-to-end speech translation that keeps voice and prosody, for cross-language communication.

Voice cloning & conversion

Clone a voice or convert speaking style for dubbing and personalized TTS.

Speaker diarization

Tell speakers apart for meeting notes and multi-speaker transcription.

Voiceprint recognition & verification

Confirm identity by voiceprint for identity checks and fraud prevention.

Wake word detection

Detect a chosen wake word for low-power voice assistant activation.

Voice activity detection VAD

Separate speech from non-speech to cut costs, segment audio and power real-time interaction.

Noise reduction & speech enhancement

Suppress noise and enhance voices for calls, recordings and meetings.

Echo cancellation & dereverberation

Remove echo and reverb for hands-free calls and meeting room audio.

Environmental sound recognition

Recognize sound events around you for security, industry and urban sensing.

Music & sound effect generation

Generate music and sound effects for content creation, games and short videos.

RAG knowledge Q&A

Retrieval-augmented generation that answers from your enterprise knowledge base with fewer hallucinations.

Document Q&A & comparison

Ask, compare and summarize across documents such as contracts, reports and papers.

Enterprise knowledge assistant

An all-in-one assistant for internal knowledge search, Q&A and writing.

Multimodal retrieval

Cross-modal search: find images or videos by text, or text by image.

Visual Q&A & scene understanding

Answer questions about images for accessibility, education and inspections.

Real-time multimodal assistant

An interactive assistant that handles speech, images and text in real time.

Tool-calling & workflow agents

Call external tools and APIs to orchestrate workflows and complete complex tasks.

GUI & browser agents

Operate graphical interfaces and browsers to finish tasks, for automation and RPA.

Coding & DevOps agents

Code generation, understanding and review, plus operations automation.

Code generation & understanding

Code completion, generation, explanation, refactoring and unit test generation.

Multi-agent orchestration for research & ops

Multiple agents working together on research, operations and analysis tasks.

Personalized ranking

Rank results by user behavior for recommendations, search and ads.

Candidate retrieval & similar items

Retrieve candidates and recommend similar items for e-commerce, content and social.

Semantic search & reranking

Semantic retrieval + reranking for more relevant search results.

Predictive maintenance & remaining useful life

Predict equipment failures and remaining useful life from multi-source data.

Streaming anomaly & change detection

Detect anomalies and sudden changes in streaming data in real time.

Time series forecasting

Forecast trends, seasonality and anomalies in time series data.

Scheduling, routing & resource optimization

Solve scheduling, routing and resource allocation optimization problems.

AMR navigation & dispatch

Mobile robot navigation, obstacle avoidance and fleet dispatch.

Drone & on-site autonomous inspection

Autonomous drone inspection, recognition and alerts.

Multi-sensor perception & fusion

Fuse cameras, radar, LiDAR and other sensors for perception.

IMU/radar/LiDAR fusion

Fuse inertial, radar and LiDAR data for localization and perception.

Visual-inertial SLAM

Visual + inertial odometry and mapping for robotics and AR.

Load & environmental forecasting

Forecast power load and environmental metrics for energy and city management.

Product links

Berserk Models · AI Tools · Aggregation Hub

GLACIER AI

Start with what you need.

Tell us which products interest you, your current systems and your goals.

shenzhongchang@glacier.fit

Product scope, configuration and activation are confirmed with you directly.