机器人柔性机械臂具有机动性好、覆盖范围大、成本低和节能等诸多优势,得到日益深入的研究和广泛的应用。但是柔性机械臂运动存在振动这一共性问题,解决振动问题是有效运用柔性机械臂高质量完成控制任务的基础和关键。 传统的柔性机械臂控制任务与目的涉及了位置与跟踪控制的残余振动抑制和跟踪控制的稳态振动抑制,解决振动问题的机理方法和效果存在一定的局限性和不足。论文研究了改进残余振动和稳态振动的控制方法,针对传统策略与方法的现有问题和连续减振需求,探索研究了解决动态的振动减除的振动避免原理和方法,并研究了机器学习的自治振动控制方法与算法。主要研究内容和创新点总结如下: (1)针对柔性机械臂的残余振动问题,研究了点到点的振动抑制控制,提出频谱激励减振方法。引入和研究关于减振的动力学模型非线性降解的局域不变性准则,研究和建立了该准则下参数变动灵敏性分析指标,给出非线性模型分段线性化计算法,数值仿真检验局域不变性。研究证明了多模态耦合和后置构型变动下一致减振的存在性,给出减振条件和多谐振零化计算模型;证明离线逆向生成时分激励的频谱激励减振计算性质,给出减振控制设计方法。研究了多模态谐振带的减振控制,研究给出带状模态减振性质,由带状减振增强频谱激励减振控制的鲁棒性。根据两连杆机构的物理模型进行控制器设计和计算,给出对象的振动控制数值仿真,检验了频谱激励减振控制的有效性。 (2)针对跟踪控制的振动问题,研究了动态平衡的减振方式,探索了刚随柔动的振动避免控制方法。通过研究材料力学和振动力学,构想了弹性体中性面变形的顺势激励方式,提出刚轴推进跟随柔杆进动的新控制原理,构建了刚随柔动、刚柔同步一体的避免振动的控制基础。研究了柔性机械臂避振控制的任务与目的、给出振动避免定义。研究了基于刚随柔动的超前和滞后型连杆中性面单侧稳恒运行的机制,以及该机制的性质和控制律,给出振动避免控制器实现。基于振动避免方法的性质研究了避振控制在形变定义域上平衡点和不变集的动平衡态。研究刚随柔动原理方法的振动控制闭环系统的稳定性。根据动态平衡的稳定性质和条件,基于Lyapunov稳定定理和LaSalle不变性定理证明了跟踪避振PD控制闭环系统关于动平衡态和正向极限点的全局渐进一致的稳定性。通过仿真验证了跟踪避振方法的有效性。 (3)针对振动避免控制,探索了增强学习递推生成控制的方式。根据带减额因子的性能指标,研究得出可含减振命令的增广状态无限时间LQT二次型;研究了跟踪减振效用的基于Markov链动态规划的最优评估Bellman方程,以及遍历性和平稳性条件下前向递归策略评价和改进计算原理,给出了最优策略的代数Riccati方程。研究了时序差分法和策略随机逼近的在线迭代算法;针对跟踪减振单样本路径的决策最优,研究了Q函数的双重功效,给出不依赖动力学知识的策略评价与改进处理;研究含输入增广状态的数值型二次型Q函数Bellman方程,给出了Q学习策略评价与改进的最优控制逼近在线前向迭代算法。通过对柔性单连杆机械臂的跟踪振动控制数值仿真,检验了在线因果递归Q学习跟踪振动控制的有效性。 关键词:柔性机械臂;振动控制;频谱激励法;振动避免;增强学习
Flexible robotic arms possess the advantages including dexterity, large working space, lower costs, and energy saving, thus gaining growing study and being widely used in many industries. However, such flexible manipulators confront with a common problem: vibration during operation. Solving the vibration problem is a basic and key point to realize effective manipulation and achieve a task with high quality. The existing control task and objective of flexible arm involves residual vibration suppression (VS) and steady vibration reduction (VR) of positioning and tracking control respectively, which are limited and deficient in principle methods and effectiveness of removing vibration. In this thesis, a new method to improve control of residual and steady vibration has been proposed, and for the demand of continuous vibration reduction and to deal with the deficiency using the conventional strategies, the principle and method of vibration avoidance has been investigated and presented to reduce dynamic vibration. The autonomous approach and computing algorithm of vibration control based on machine learning is further studied and explored. The main study and the innovative work are summarized as follows: 1.For residual vibration of a flexible arm, the VS of point-to-point motion is studied and the impulse spectrum is proposed. The local invariance criterion for the reduction of nonlinearity of the dynamic model is introduced and studied for VR processing, and the sensitivity index associated with the criterion computation is established, with the segmented LTI pieces for control approach under the model given for computing. The calculation to verify the criterion is conducted. The existence of uniform VR of coupled multimodes and configuration variation is proved in the study, and the VR conditions and the annihilation factors are provided. The input inverse approach of VR by the impulse spectrum based on the time slicing ignition is developed and the control design procedures are discussed. The control of VR for the multimode band is investigated and the property of the impulse spectrum is analyzed for preparing the VR control over the band-wide multimodes enhancing the control robustness. The validity of the impulse spectrum control is demonstrated by numerical simulations. 2.For VR of tracking control, the concept to immune vibration of a dynamic process has been investigated and the method with the rigid-to-follow-flexible (RTFF) pace is proposed for vibration avoidance (VA). The pace-following excitation is conceived by studying mechanics of materials and vibration for deformation of the neutral surface of an elastic link, and the new principle of rigid-propelling-to-follow-flexible-precession is presented to build up the RTFF movement as the basis to constitute rigid-flexible synchronization for VA. The operational task and objective of a flexible arm for VA is proposed, with the vibration avoidance defined by analyzing the avoidance behavior and essence. The mechanism of single-sided steady run of the neutral surface is examined based on the RTFF for either lead or hysteretic deformation. The property of the mechanism and the regulating law using the mechanism are analyzed and studied to provide the implementation of the controller accordingly. In regard to VA, the dynamic equilibrium of an equilibrium point and an invariant set in the defined domain of deformation is analysed in terms of the properties of the new method. The stability of the closed-loop system with the vibration control using the RTFF has been examined. Under the stable condition and property analyzed, the closed loop system with the PD controller for VA tracking is proven globally uniformly asymptotically stable for a dynamic equilibrium set and its positive limit point based on the Lyapunov and LaSalle's arguments. The numerical simulations of a single link flexible manipulator validate the proposed VA control method. 3.With regard to vibration control, an approach to generate recursively a VA controller with reinforcement learning is studied. In line with the performance index derated by a discount factor, the quadratic form of infinite horizon LQT for the augmented state with possible VR commands is examined and produced; the study on the Bellman equation of optimal evaluation based on the dynamic programming of Markov chain is carried out, with the utility of tracking and VR; the computing mechanism on the policy evaluation and improvement of forward recursion on the condition of ergod--icity and stationarity is studied as well, and yielding the algebraic Riccati equation is provided for the optimal policy. The online iteration algorithm based on the temporal-difference learning and the stochastic approximation of policy is studied; for policy optimality of tracking VR along a single sample trajectory, the dual efficacy of the Q function has been examined and developed, yielding the processing to convert the system knowledge based into the data based policy evaluation and improvement; the Bellman equation of the numeric Q function with possible augmented states including an input is studied, and the online forward iteration algorithm for optimal control approaching using the policy evaluation and improvement by Q learning is presented. A flexible manipulator of single link is used for simulations that validate the effectiveness of the proposed Q learning optimal tracking vibration control. KEY WORDS :flexible manipulator; vibration control; impulse spectrum method; vibration avoidance; reinforcement learning